← uncontaminatable benchmarks

jcode bench

Optimize production-grade primitives. Exhaustive correctness is required. Every verified improvement counts, and runs have no time limit.

How a run works

Each task gives the agent a working implementation, an exhaustive verifier, and a published deterministic cost model. The agent edits the code and submits it for grading as often as needed. Each grade takes seconds, and the benchmark records every new best result. Correctness on every possible input is the gate. Among passing submissions, the score measures doublings of cost reduction relative to the starting implementation.

Tasks

Three tasks, chosen for fast grading and substantial optimization headroom. A better result raises the known frontier for every later run.

float-print: shortest round-trip float to decimal

Convert a float32 to the shortest decimal string that parses back to the same bits. This remains an active research area, with algorithms including Grisu, Ryū, and Dragonbox. The verifier checks all 232 float values.

json-unescape: decode JSON string escapes

Decode a performance-critical part of JSON parsing. Escape density varies across inputs, leaving room for strategies beyond current SIMD implementations. The verifier checks every bounded-length sequence.

utf16-transcode: UTF-16 to UTF-8

Convert UTF-16 data used by JavaScript and Windows APIs into UTF-8. Mixed-width branching leaves optimization room. The verifier checks all code units, all surrogate pairs, and boundary-stratified combinations.

Frontier model comparison →

Leaderboards: every harness and model, by task

Every run below is one solo agent that passed the official final correctness grade. Reasoning effort is shown where a campaign includes multiple levels. Each chart plots the running best score against active agent time. The two default lines are the original 2026-07-05/06 jcode and Claude Code head-to-head runs. The buttons add or remove any other harness, model, or reasoning effort.

These charts rank task by task across harnesses. For the model-only ranking with a single combined score, see the model comparison: six frontier models through the same pinned Jcode harness, aggregated by geometric-mean doublings.

json-unescape

+0.5 +1.0 +1.5 +2.0 +2.5 +3.0 0m 30m 60m 90m 120m 150m jcode +2.69 claude code +2.09
#harnessmodelbest scorespeedupgradesactive timedate
1jcodeClaude Fable 5+2.83907.2x782.3h2026-07-19
2jcodeClaude Opus 4.8+2.68896.4x1632.6h2026-07-05
3Codex CLIGPT-5.6 Sol+2.42285.4x2120m2026-07-10
4jcodeGPT-5.6 Sol+2.30554.9x2619m2026-07-19
5OpenCodeGPT-5.6 Sol+2.18944.6x2528m2026-07-17
6Claude CodeClaude Opus 4.8+2.09204.3x2658m2026-07-05
7OpenCodeClaude Opus 4.8+1.99914.0x8343m2026-07-18
8jcodeGPT-5.5+1.98414.0x339m2026-07-19
9jcodeGPT-5.4+1.70193.3x206m2026-07-19
10jcodeClaude Sonnet 5+1.2270one degenerate sample excluded†2.3x4549m2026-07-19
jcodeClaude Opus 5 (high)+3.04538.3x333.8h2026-08-24
Claude CodeClaude Opus 5 (high)+3.01938.1x791.1h2026-08-24
Claude CodeClaude Opus 5 (low)+2.95657.8x3142m2026-08-24
jcodeClaude Opus 5 (low)+2.49115.6x6233m2026-08-24
jcodeOx Alpha (low)+2.35085.1x221.8h2026-08-24
jcodeOx Alpha (high)+1.52422.9x540m2026-08-24
jcodeOx Alpha (max)+1.96863.9x361.8h2026-08-24

float-print

+1 +2 +3 +4 +5 +6 +7 +8 +9 0h 1h 2h 3h 4h 5h 6h 7h 8h 9h 10h 11h jcode +8.64claude code +7.17
#harnessmodelbest scorespeedupgradesactive timedate
1jcodeClaude Fable 5+12.00864120x366.5h2026-07-19
2jcodeClaude Opus 4.8+8.6385399x2810.6h2026-07-06
3jcodeGPT-5.6 Sol+7.8107225x2348m2026-07-19
4Codex CLIGPT-5.6 Sol+7.4165171x3042m2026-07-10
5OpenCodeGPT-5.6 Sol+7.2181149x1436m2026-07-17
6OpenCodeClaude Opus 4.8+7.2077148x3395m2026-07-18
7jcodeGPT-5.5+7.2042147x2352m2026-07-19
8Claude CodeClaude Opus 4.8+7.1692144x2486m2026-07-06
9jcodeGPT-5.4+7.0336131x948m2026-07-19
10jcodeClaude Sonnet 5+6.8199113x382.4h2026-07-19
Claude CodeClaude Opus 5 (high)+10.16981151.9x422.8h2026-08-24
jcodeClaude Opus 5 (high)+8.5787382.3x924.1h2026-08-24
Claude CodeClaude Opus 5 (low)+8.4412347.6x441.0h2026-08-24
jcodeClaude Opus 5 (low)+7.7596216.7x591.8h2026-08-24
jcodeOx Alpha (low)+6.9037119.7x420m2026-08-24
jcodeOx Alpha (high)+7.0329131.0x102.2h2026-08-25
jcodeOx Alpha (max)+7.5686189.8x201.8h2026-08-25

utf16-transcode

+1 +2 +3 +4 0h 1h 2h 3h 4h 5h 6h 7h 8h 9h 10h 11h jcode +3.28claude code +2.39
#harnessmodelbest scorespeedupgradesactive timedate
1jcodeClaude Opus 4.8+3.27979.7x11610.5h2026-07-06
2jcodeClaude Fable 5+2.55155.9x794m2026-07-19
3Claude CodeClaude Opus 4.8+2.39035.2x55102m2026-07-06
4jcodeGPT-5.6 Sol+2.11424.3x1812m2026-07-19
5OpenCodeClaude Opus 4.8+1.86383.6x4232m2026-07-18
6Codex CLIGPT-5.6 Sol+1.83273.6x2417m2026-07-10
7OpenCodeGPT-5.6 Sol+1.50842.8x1925m2026-07-17
8jcodeGPT-5.5+1.34112.5x177m2026-07-19
9jcodeClaude Sonnet 5+1.25612.4x3790m2026-07-19
10jcodeGPT-5.4+1.03312.0x228m2026-07-19
Claude CodeClaude Opus 5 (high)+3.361510.3x501.7h2026-08-24
jcodeClaude Opus 5 (high)+2.91927.6x82.6h2026-08-24
Claude CodeClaude Opus 5 (low)+2.70026.5x2643m2026-08-24
jcodeClaude Opus 5 (low)+2.40155.3x8753m2026-08-24
jcodeOx Alpha (low)+0.00001.0x21m2026-08-24
jcodeOx Alpha (high)+1.87283.7x61.6h2026-08-24
jcodeOx Alpha (max)+1.61543.1x133.0h2026-08-24

Best is the highest grade sampled during the run, so it can sit slightly above the official final score, and for float-print not every sample is a full 232 gate. Runs were recorded on different dates with different harness versions, and each cell is one run. †The Claude Sonnet 5 json-unescape run logged a single +14.30 grade from a degenerate random corpus; the surrounding grades sat near +0.93 and no later grade came close, so it is excluded here and documented in the raw curve data.

All task specs, graders, verifiers, and given implementations are public at github.com/1jehuang/jcode-bench. There is no hidden test set. Read about the benchmark class at /bench.