AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Paper • 2608.26623 • Published 10 days ago • 19
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments Paper • 2608.24804 • Published 12 days ago • 39
Recursive Synthesis for Long-Horizon Terminal Tasks Paper • 2608.05466 • Published 30 days ago • 252
💫StarShell Collection Resources for paper "Terminal Agents Suffice for Enterprise Automation". Repo: https://github.com/servicenow/starshell • 7 items • Updated about 1 month ago • 14
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 72
Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published Jun 27 • 151
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 124
view article Article Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech ServiceNow-AI • Jun 9 • 46
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models Paper • 2510.04618 • Published Oct 6, 2025 • 134
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents Paper • 2605.13841 • Published May 13 • 78
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics Paper • 2605.12178 • Published May 12 • 66
view article Article Vision Language Models (Better, faster, stronger) +3 merve, sergiopaniego, ariG23498, pcuenq, andito • May 12, 2025 • 617
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Paper • 2604.06628 • Published Apr 8 • 330
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning Paper • 2604.02007 • Published Apr 2 • 14