view article Article Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv sergiopaniego • 7 days ago • 22
view article Article Can you train a model on Simon Willison's deeply unscientific pelican benchmark? sergiopaniego • 13 days ago • 2
view article Article Shipping huggingface_hub every week with AI, open tools, and a human in the loop Wauplin, celinah • Jun 23 • 20
ClaimExtractor-2605 Collection Extract claims and intents from conversations • 7 items • Updated Jun 14 • 8
view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 137
view article Article Multimodal Embedding & Reranker Models with Sentence Transformers tomaarsen • Apr 9 • 70
LFM2 2.6B Mr. Tic Tac Toe ❌ ⭕ Collection Dataset and models for transforming LFM2 2.6B into a Tic Tac Toe master using RL Environments. Free course: https://t.ly/4jIFq • 8 items • Updated Apr 8 • 2
view article Article TRL v1.0: Post-Training Library Built to Move with the Field +2 qgallouedec, stevhliu, pcuenq, sergiopaniego • Mar 31 • 58
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation Paper • 2602.17316 • Published Feb 19 • 2
Zagreus - Nesso fine tuned Collection The collection contains three bilingual English/Italian SLMs post-trained on Zagreus-0.4B-ita: instruct, agentic, and a fully open-source • 3 items • Updated Mar 4 • 3
Zagreus 0.4B Collection The Zagreus-0.4B collection contains four bilingual English + Romance language foundational SLMs (~400M parameters) trained from scratch • 4 items • Updated Mar 4 • 7
view article Article Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries +7 aminediroHF, qgallouedec, kashif, lewtun, edbeeching, albertvillanova, nouamanetazi, lvwerra, sergiopaniego • Mar 10 • 174
Qwen3.5-text-only Collection Text-only versions of Qwen-3.5 without the vision encoders for a smaller memory and storage footprint. • 4 items • Updated Jun 5 • 15