Sublate.
Every change must earn its place.
Sublate turns your scheduling, packing or routing problem into a scored benchmark. Then it improves your algorithm 24/7, and promotes only the changes that beat your baseline on held-out and stress tests.
● Director Claude sets the next 6-hour search plan ● Propose candidate code changes, each in a sandbox ● Evaluate official scorer → holdout → stress [promoted] better on all three gates [blocked] better on dev, worse on holdout ● Review Claude + a second reviewer: PASS ● Ship improved code + evidence report >
Your algorithm gets better by being challenged.
Your current algorithm
The champion that runs today. It is the bar that every change must clear.
Challengers
The loop writes candidate changes around the clock. Each one attacks a weak spot of the champion.
A verified new champion
Only a change that wins on held-out and stress tests is promoted. It then becomes the next thesis.
The name: to sublate (Hegel's Aufheben) is to cancel, keep and lift up at the same time.
From a problem with a score to a verified gain.
Formulate
We turn the problem, the objective and the constraints into an evaluator that uses your official scoring.
Baseline
Your current algorithm, or a safe heuristic, becomes the champion to beat.
Loop
Fast models propose code changes in bulk. Every candidate runs in a sandbox. Claude sets the search direction every six hours.
Gate
A change is promoted only if it beats the champion on held-out and stress instances and passes a review.
Deliver
You get the improved code and an evidence report: what changed, by how much, and on which instances.
Optimization Grand Challenge 2026
Shipyard block placement and scheduling with bay and crane constraints.
Improvement of the official objective sum on public training instances P1–P6, from the first valid solution of the loop to its final adopted solution.
Regression leaks across more than 250 recorded instance runs.
Errors or catastrophic outputs in a 94-case full evaluation.
Instances proven optimal (gap 0 against a lower bound).
Measured on the public training instances of the organizer. Final scoring used a private test set, so these numbers are not a ranking. The same system ran a second track at the same time: physics-simulated random palletizing.
Safe enough to leave on overnight.
Evaluators decide
Only local evaluators that wrap the official scorer judge a candidate. Models propose, steer and review. They never score.
No regressions
A new champion must win on held-out and stress sets. Candidates that only fit the dev set are blocked automatically.
Sandboxed candidates
Every candidate program runs in a sandbox, apart from the champion and the evaluators.
Separated roles
Proposing, judging, steering and verifying are separate components, so one failure cannot approve itself.
Budget caps
Daily cost caps and per-call limits stop a runaway loop before it burns the budget.
Watchdog and digests
A watchdog checks heartbeat, disk, memory and budget. You get a short digest twice a day.
Any problem with a score.
If you can score a solution, the loop can improve the algorithm that produces it.
Claude steers. Evaluators decide.
Claude (Opus, run headless through Claude Code) is the director that sets the search direction four times a day, and one of the reviewers that gate each promotion. Fast, low-cost models generate candidate code in bulk.
Have a problem with a score? Let us run a pilot.
Send us the problem, the objective and a sample of instances. We reply with a pilot plan.