view article Article TutorMoments: Do AI tutors know when to help and when to hold back? allenai • 1 day ago • 17
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published 18 days ago • 14
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Paper • 2608.03994 • Published 5 days ago • 6
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 9 days ago • 94
Instella-MoE ✨ Collection Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs. • 6 items • Updated 12 days ago • 17
view article Article NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 24 days ago • 58
Nemotron 3 Embed Collection Open embedding models for enterprise RAG, agentic retrieval, code search, and agent memory. • 3 items • Updated 23 days ago • 35
view article Article Distillation in 2026 (so far): which frontier models use it and how sergiopaniego • Jul 8 • 19
SWE-FastContext Collection A family of code-search models powering the Explore subagent for coding agents.(It will be made public later) • 3 items • Updated Jun 30 • 18
Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated 23 days ago • 191
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis Paper • 2605.30434 • Published May 28 • 23
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 124
Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning Paper • 2605.28424 • Published May 27 • 32