Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
-
Updated
Jun 17, 2026 - Python
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
[NeurIPS 2025] AGI-Elo: How Far Are We From Mastering A Task?
Benchmark methodology, task sets, and evaluation results for RA²R
A reproducible LiveCodeBench evaluation scaffold for controlled studies of code-reasoning SFT on Qwen2.5-1.5B.
LoRA fine-tuning evaluation pipeline for code generation models on LiveCodeBench benchmark
NeoSmith Maestro — frontier-level coding accuracy from a bouquet of small models at ~1/30th the cost. Technical report + reproducible evidence (LiveCodeBench v6 92.2%, SWE-bench Pro 74.2/88.9).
Self-Improving Agent for Code Generation
“U.S. Provisional Patent Application No. 63/978,753, filed February 9, 2026” Aleph Generalizable Intelligence Convergence Engine (AGICE) - Evidence-Native, Policy-Governed Reasoning with Rollback and Geometrical Multi-Dimensional Transformation (GMDT) via the Complex Operator (−i)
Executable-code benchmark and verifier-guided post-training sandbox in the L20 project family.
Interactive benchmark for evaluating LLMs on clarifying ambiguous code requirements. Paper: arXiv:2607.00711
Add a description, image, and links to the livecodebench topic page so that developers can more easily learn about it.
To associate your repository with the livecodebench topic, visit your repo's landing page and select "manage topics."