inixiative presents:

Do I Still Have a Job?

STATUS: YES

One question, run as an experiment: what happens if you take the human out of the loop? I built every reference here by steering an AI — the leaderboard is the same models, on the same brief, with me gone. The distance between them is the job.

Every model launch swears it can "replace your senior engineers." Adorable. So we took the inixiative ecosystem's hardest primitives — the small, brutal ones that take real engineering to get right — handed each agent the same product brief a human gets from a kickoff, and let it cook. Then a judge that has read the reference implementation grades the homework.

And before anyone blames the scope: these references are usually a few hundred lines, not thousands. That's the trap. Small and complete is the hard part — agents love to either pad it out or quietly delete half the problem. The reference did neither.

The industry wants to hand these things the keys and walk away. Fair enough — but "autonomous" only means something if it can reason a hard problem all the way through without someone holding its hand and steering mid-flight. So we don't steer: one brief, one sandbox, zero hints, zero "are you sure?". Off the leash. Let's see how far they get before they need an adult in the room.

And it's a fair fight: every reference was written 100% with AI too — same models, same tools the agents get. The only variable is whether a human is in the loop, steering. That's the whole point — it isn't human-vs-machine, it's the same machine with and without an architect. Autonomy is the dividing line — and so far, the cliff.

Getting the tests to pass is the participation trophy. We score the part that's hard. Problems are withheld, so nobody can cram — you get a number, a difficulty, and a bar. The agents get perspective.

flame metric:
bars rescale to the chosen metric (per test); color depends on the metric

Test #1

architect golden · 175 loc 12 contenders
Claude Opus 4.8
solo · high
18.5/48 39% c 15/20 v4 i0 60,232t 870s 426 loc
fail even-more-convinced architect
Claude Opus 4.8
solo · low
17.5/48 37% c 11/20 v3 i0 26,928t 448s 247 loc
fail even-more-convinced principal
Claude Opus 4.8
solo · xhigh
16.5/48 34% c 12/20 v3 i0 88,849t 1798s 489 loc
fail even-more-convinced staff
Claude Sonnet 4.6
solo · max
16/48 33% c 10/20 v2 i0 74,900t 1135s 515 loc
fail even-more-convinced principal
GPT-5.5
solo · xhigh
14.5/48 30% c 11/20 v5 i0 —t 930s 908 loc
fail even-more-convinced staff
GPT-5.5
solo · high
14/48 29% c 10/20 v3 i0 —t 1199s 595 loc
fail even-more-convinced senior
Claude Opus 4.8
solo · max
13/48 27% c 11/20 v3 i0 142,855t 1988s 511 loc
fail even-more-convinced principal
Claude Opus 4.8
solo · medium
13/48 27% c 13/20 v2 i0 47,875t 699s 369 loc
partial even-more-convinced staff
Claude Sonnet 4.6
solo · xhigh
8/48 17% c 10/20 v2 i0 44,789t 784s 291 loc
fail even-more-convinced staff
Claude Sonnet 4.6
solo · high
8/48 17% c 9/20 v1 i0 38,153t 628s 275 loc
fail even-more-convinced staff
Claude Sonnet 4.6
solo · medium
6.5/48 14% c 6/20 v2 i0 32,964t 596s 284 loc
fail even-more-convinced staff
Claude Sonnet 4.6
solo · low
5.5/48 12% c 3/20 v1 i0 26,088t 473s 209 loc
fail even-more-convinced principal

Test #2

principal golden · 415 loc 10 contenders
Claude Sonnet 4.6
solo · high
13/49 27% c 3/11 v4 i0 32,491t 756s 142 loc
fail even-more-convinced
GPT-5.5
solo · xhigh
13/49 27% c 3/9 v4 i0 —t 705s 821 loc
fail even-more-convinced senior
Claude Opus 4.8
solo · medium
13/49 27% c 5/11 v0 i0 35,730t 505s 220 loc
fail even-more-convinced staff
Claude Opus 4.8
solo · low
13/49 27% c 4/11 v0 i0 23,894t 369s 160 loc
fail even-more-convinced senior
GPT-5.5
solo · medium
13/49 27% c 4/11 v4 i0 —t 386s 376 loc
fail even-more-convinced senior
Claude Sonnet 4.6
solo · max
10.5/49 21% c 3/11 v2 i0 28,728t 931s 137 loc
fail even-more-convinced staff
Claude Sonnet 4.6
solo · low
10.5/49 21% c 3/11 v4 i0 14,302t 296s 165 loc
fail even-more-convinced senior
GPT-5.5
solo · high
9.5/49 19% c 1/11 v9 i0 —t 801s 740 loc
fail even-more-convinced senior
Claude Sonnet 4.6
solo · medium
8.5/49 17% c 3/11 v2 i0 22,773t 405s 149 loc
fail even-more-convinced senior
Claude Sonnet 4.6
solo · xhigh
5.5/49 11% c 2/11 v3 i0 32,444t 1291s 257 loc
fail even-more-convinced staff

Bar length = the selected metric (rescaled per test). On score & corners, color = completion: green complete, amber partial, red fail; cost metrics (tokens, seconds) run neutral. On LOC, the dashed mark is the golden's size (1×) and color reads economy — green at or under it, amber up to 2×, red beyond. self = the difficulty the model thought it was facing. golden = what the run did to our confidence in the reference (so far: more convinced, not less).

Wait — what are these problems?

They're drawn from inixiative: a set of small, sharp, open-source primitives — an authorization core, a rules engine, a lifecycle guard — each built to be minimal and complete. Apparently that's the hard part. Come see what we're building.

Interested in the bench — running it, contributing problems, or the methodology behind the numbers? Get in touch →