Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
RL+LLM Wiki
community
Activity Feed
Follow
26
AI & ML interests
None defined yet.
Recent Activity
lvwerra
new
activity
about 19 hours ago
rl-llm-wiki/knowledge-base:
source: url:pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2 — DeepScaleR 1.5B iterative-context-lengthening RL reproduction (open-repro)
lvwerra
new
activity
1 day ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/summertime-outlook-o3s-novelty-coming — o3 novelty / RLVR search (speculation)
lvwerra
updated
a bucket
1 day ago
rl-llm-wiki/rl-main-bucket
View all activity
Team members
14
rl-llm-wiki
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Articles
lvwerra
in
rl-llm-wiki/knowledge-base
about 19 hours ago
source: url:pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2 — DeepScaleR 1.5B iterative-context-lengthening RL reproduction (open-repro)
3
#773 opened 4 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
1 day ago
source: url:interconnects.ai/p/summertime-outlook-o3s-novelty-coming — o3 novelty / RLVR search (speculation)
4
#728 opened 6 days ago by
lvwerra
lvwerra
updated
a bucket
1 day ago
rl-llm-wiki/rl-main-bucket
319 MB
lvwerra
in
rl-llm-wiki/knowledge-base
1 day ago
source: url:dwarkesh.com/p/john-schulman — Dwarkesh x John Schulman: pre-o1 reasoning/RL insider framing (transcript)
4
#775 opened 4 days ago by
lvwerra
source: url:nishtahir.com/notes-on-the-phi-4-reasoning-technical-paper — Phi-4-reasoning recipe reconstruction (Microsoft, secondary)
4
#765 opened 4 days ago by
lvwerra
source: url:newsletter.semianalysis.com/p/deepseek-debates — DeepSeek true training cost / $6M myth (speculation)
3
#740 opened 4 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
2 days ago
source: url:cnblogs.com/theseventhson/p/18725466 — Kimi k1.5 long-CoT RL reconstruction (CN, speculation)
4
#763 opened 4 days ago by
lvwerra
source: url:j-qi.medium.com/what-does-rl-improve-when-it-improves-llm-reasoning-2befa16c56e8 — RL = distribution shaping / distillation-vs-RL debate (speculation)
4
#772 opened 4 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
3 days ago
source: url:dwarkesh.com/p/dario-amodei-2 — Dwarkesh x Dario Amodei: RL scales log-linear like pretraining (transcript)
3
#778 opened 4 days ago by
lvwerra
source: url:dwarkesh.com/p/demis-hassabis — Dwarkesh x Demis Hassabis: AlphaZero-atop-LLMs, pre-o1 (transcript)
3
#779 opened 4 days ago by
lvwerra
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking — RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
3
#784 opened 4 days ago by
lvwerra
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying — OpenAI's many RLs / agentic (speculation)
4
#734 opened 6 days ago by
lvwerra
source: url:mechanize.work/blog/the-upcoming-gpt-3-moment-for-rl — RL environments as scaling substrate / GPT-3-moment thesis (forecast)
3
#785 opened 4 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
4 days ago
source: url:interconnects.ai/p/opening-the-black-box-of-character — Character training pipeline (speculation, paid/partial)
3
#732 opened 6 days ago by
lvwerra
source: url:pillumina.github.io/posts/aiinfra/02-slime — Zhipu GLM slime RL-infra source reconstruction (CN, speculation)
3
#764 opened 4 days ago by
lvwerra
source: url:interconnects.ai/p/thinking-searching-and-acting — Thinking/Searching/Acting (agentic speculation)
3
#731 opened 6 days ago by
lvwerra
source: url:yuanchaofa.com/post/kimi-k2-5-reading-notes — Kimi K2.5 PARL parallel-agent RL deep-read (CN, speculation)
3
#744 opened 4 days ago by
lvwerra
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
4
#739 opened 4 days ago by
lvwerra
source: url:cnblogs.com/volcengine-developer/articles/19070102 — veRL+ReTool tool-use RL reproduction / env+loss plumbing (CN, speculation)
3
#770 opened 4 days ago by
lvwerra
source: url:hkust-nlp.notion.site/simplerl-reason — SimpleRL-Zero 7B/8K R1-Zero reproduction / easy-to-hard (open-repro)
3
#774 opened 4 days ago by
lvwerra
Load more