MassAlloc Attention: Let Attention Allocate Its Own Compute Paper • 2609.32712 • Published 3 days ago • 21
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations Paper • 2609.30222 • Published 5 days ago • 8
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 4 days ago • 118
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure Paper • 2609.26550 • Published 7 days ago • 42
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup Paper • 2609.15126 • Published 15 days ago • 12
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 12 days ago • 57
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 12 days ago • 53
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 12 days ago • 189
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 12 days ago • 75
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 15 days ago • 51
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 21 days ago • 81
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • Jul 10 • 50
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 19 days ago • 702
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements Paper • 2509.01809 • Published 21 days ago • 4
view article Article Bringing Nunchaku 4-bit Diffusion Inference to Diffusers rootonchair, sayakpaul • Jul 23 • 69