LoRA Speedrun: Public Leaderboard for Fine-Tuning Speed
20 Jul 2026
A new public benchmark is turning LLM fine-tuning into a timed sport. Called LoRA Speedrun, it's a wall-clock leaderboard that pits fine-tuning techniques against each other on a frozen task and frozen hardware — and the current record for tuning Qwen2.5-1.5B to hit 57%+ accuracy on GSM8K sits at just 6 minutes 5 seconds.
How it works
The rules are deliberately narrow, which is the point:
- Task: Fine-tune Qwen2.5-1.5B on GSM8K to reach an accuracy threshold of 57% or higher.
- Hardware: A single L40S GPU — no multi-GPU setups allowed.
- Data: Training is restricted to the GSM8K train split only. Teacher models, synthetic data augmentation, and outside data sources are banned.
- Verification: Every claimed record must be independently re-run three times with fresh seeds before it counts.
For local iteration, participants need a minimum of 24 GB of GPU memory. Notably, the report indicates participants can train on as few as 1,000 examples in 90 seconds if they can still clear the accuracy bar — suggesting speed and data efficiency are as valuable as raw compute.
The current record
The leaderboard's current best-known time — 6 minutes 5 seconds — was set by a participant using sequence packing and completion-only loss masking over 2 epochs, reportedly achieving roughly 2x faster results at higher accuracy than the baseline. The report does not specify how many prior records existed before this one or how long the leaderboard has been running.
Open ground for optimization
According to the report, a long list of techniques remains unclaimed on the leaderboard, including:
- 1-epoch aggressive learning-rate schedules
- Data pruning
- Block-diagonal/varlen packing attention
- QLoRA NF4 vs. bf16 tradeoffs
- rsLoRA/DoRA/PiSSA initialization
- LoRA+
- NEFTune noise injection
- Curriculum ordering
- Rank/placement search
- torch.compile, Unsloth kernels, Liger kernels, fused cross-entropy
- Smarter warmup schedules
This open list functions as an implicit invitation for anyone experimenting with fine-tuning efficiency to try their hand at the record.
Verification, licensing, and cost
All verification reports and security reviews are posted publicly on pull requests, and the project is MIT-licensed, with records, reports, and write-ups publicly available. Modal's free monthly compute credits reportedly cover full verification runs on the L40S sandbox, which could lower the cost barrier for participants who don't have dedicated GPU access.
The report does not detail who performs the security reviews beyond noting they're posted on pull requests, nor does it mention any prizes or formal incentives tied to setting a record.
Why founders should care
For early-stage founders working with LLM fine-tuning, this leaderboard is likely more useful as a benchmarking reference than a production blueprint. A few considerations:
- Because the setup is fixed to a single L40S GPU and a narrow dataset, techniques that win here may not directly generalize to multi-GPU or large-scale production training — founders should treat results as directional, not definitive.
- The emphasis on three independently verified re-runs suggests a reproducibility standard that founders could reasonably borrow for their own internal benchmarking of fine-tuning pipelines.
- Given the MIT license and public verification trail, founders may be able to study or reuse techniques from top entries with relatively low legal risk, though the report gives no specifics on the code's production-readiness.
- The requirement for reliable repeated compute access to complete three verified re-runs could mean participation — and by extension, visibility into winning techniques — skews toward teams or individuals with steady GPU access, potentially limiting how representative the leaderboard is of broader practitioner capability.
Overall, LoRA Speedrun looks like a niche but potentially informative signal for founders evaluating fine-tuning efficiency techniques, provided they keep its single-GPU, single-dataset constraints in mind before extrapolating results to their own stacks.