On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification Paper • 2508.05629 • Published Aug 7, 2025 • 191
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated 2 days ago • 78
RUT-Bench Collection Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions". • 2 items • Updated Jun 4 • 1
👤 Implicit Personalization in Language Models Collection Works on detecting, attributing and controlling implicit personalization in language models • 29 items • Updated Mar 20 • 4
Finetuning LLMs for Human Behavior Prediction in Social Science Experiments Paper • 2509.05830 • Published Sep 6, 2025 • 1
Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation Paper • 2508.18142 • Published Aug 25, 2025 • 1
Learning from Language Feedback via Variational Policy Distillation Paper • 2605.15113 • Published May 18 • 13
Flipping the Dialogue: Training and Evaluating User Language Models Paper • 2510.06552 • Published Oct 8, 2025 • 2
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents Paper • 2604.26752 • Published Apr 29 • 114
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning Paper • 2605.19461 • Published May 19 • 2