Limes Research Programs is the public research agenda for Limes Labs.
The purpose of this repository is discipline, not hype. Limes Labs is building falsifiable experiments, reusable evaluation infrastructure, and public research protocols before making frontier-capability claims.
Limes Labs is not claiming to be a frontier lab today. The current objective is to build the habits and artifacts that make serious research possible:
- preregistered experiments
- clear baselines and controls
- negative-result logs
- benchmark designs that resist easy saturation
- promotion gates before scale-up
- reusable templates for public contributors
Public claims should stay proportional to public evidence. A result can be interesting before it is decisive; it should be labeled that way.
Start with the agenda overview:
- Public research agenda
- Long-horizon RL
- Small-model efficiency
- Optimizer research
- Benchmark design
- Agentic LLM challenges
- AutoResearch
- European applied evals
These templates are meant to be copied into experiment repos before results are known:
The public extraction map turns private research themes into sanitized public work items without publishing private code, raw logs, secrets, or unreviewed claims:
- limes-autoresearch
- limes-dataforge
- limes-kernelforge
- limes-nanogpt
- limes-parameter-golf
- eurobench
- model-card-template
- limes-constitution
A good contribution makes one question easier to test. It does not need to make the lab sound bigger than it is.
Prefer small, replayable work:
- State the question.
- Lock the metric, baselines, budget, and selection rule.
- Run the smallest meaningful experiment.
- Publish the artifact and limitations.
- Promote only if the gate was declared before looking at the result.
This repository is an initial public scaffold. It is useful when it creates issues, reports, templates, and experiments that other Limes repositories can run and review.