Data and models for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards".
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
WildReward: Learning Reward Models from In-the-Wild Human Interactions
DeepPrune: Parallel Scaling without Inter-trace Redundancy
models 81
THU-KEG/DeepDive-30B-A3B-C-GRPO
31B • Updated • 5
THU-KEG/DeepDive-4B-C-GRPO
4B • Updated • 15
THU-KEG/DeepDive-30B-A3B-SFT
31B • Updated • 2
THU-KEG/DeepDive-4B-SFT
4B • Updated • 48
THU-KEG/WildReward-8B
Text Classification • 8B • Updated • 13 • 3
THU-KEG/WildReward-4B
Text Classification • 4B • Updated • 24 • 4
THU-KEG/LLaDA-8B-BGPO-sudoku
Reinforcement Learning • 8B • Updated • 2 • 1
THU-KEG/LLaDA-8B-BGPO-countdown
Reinforcement Learning • 8B • Updated • 129 • 1
THU-KEG/LLaDA-8B-BGPO-code
Reinforcement Learning • 8B • Updated • 5 • 1
THU-KEG/LLaDA-8B-BGPO-math
Reinforcement Learning • 8B • Updated • 2 • 1
datasets 22
THU-KEG/WildFB
Updated • 36 • 2
THU-KEG/CaRR-DeepDive
Preview • Updated • 50 • 1
THU-KEG/AgentIF
Viewer • Updated • 707 • 145 • 7
THU-KEG/DeepPrune
Preview • Updated • 8 • 2
THU-KEG/LinguaLens-Data
Viewer • Updated • 7.25k • 5 • 2
THU-KEG/RM-Bench
Viewer • Updated • 1.33k • 1.69k • 9
THU-KEG/LongWriter-Zero-RLData
Viewer • Updated • 8.61k • 33 • 21
THU-KEG/Arena-Write
Viewer • Updated • 595 • 19 • 5
THU-KEG/LongStory
Viewer • Updated • 5.28k • 14 • 3
THU-KEG/IF-Verifier-Data
Viewer • Updated • 131k • 64 • 4