Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub Paper • 2610.11169 • Published 3 days ago • 1 • 2
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks Paper • 2610.11794 • Published 3 days ago • 28 • 5
A self-learning scientific agent for X-ray diffraction Paper • 2610.07862 • Published 5 days ago • 17 • 6
Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution Paper • 2610.00417 • Published 11 days ago • 2 • 4
UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents Paper • 2610.05622 • Published 7 days ago • 22 • 3
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 10 days ago • 60 • 3
Safety of Latent Communication in Multi-Agent Systems Paper • 2609.39788 • Published 11 days ago • 9 • 2
SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents Paper • 2609.34518 • Published 13 days ago • 4 • 3
Nereus: Adaptive Parallelism for LLM Post-Training Paper • 2609.34645 • Published 13 days ago • 20 • 3
Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery Paper • 2609.27980 • Published 18 days ago • 6 • 3
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 19 days ago • 164 • 6
Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion Paper • 2609.24220 • Published 20 days ago • 68 • 4
Retention-Constrained Post-Training Quantization of Cellpose-SAM for Stem Cell Microscopy Paper • 2609.21038 • Published 24 days ago • 5 • 4
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 24 days ago • 44 • 3
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling Paper • 2609.19499 • Published 25 days ago • 37 • 4
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 24 days ago • 115 • 5
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Paper • 2609.18063 • Published 25 days ago • 26 • 4
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems Paper • 2609.17320 • Published 26 days ago • 4 • 3
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus Paper • 2609.15504 • Published 27 days ago • 37 • 5
Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents Paper • 2609.05824 • Published Sep 5 • 7 • 3