Laguna: thinking may be regime-dependent, not simply net-negative (needs review + independent test)#10
Laguna: thinking may be regime-dependent, not simply net-negative (needs review + independent test)#10TheTom wants to merge 2 commits into
Conversation
…NDENT TEST) Proposes that thinking is regime-dependent rather than simply net-negative: ON wins single-turn codegen (+2.64 pts, flakiness halves 24->11, HumanEval+ n=492), OFF wins behavioral/integrity and long-agentic. Also adds ~12k budget guidance and 'cap-hits are degeneration loops, not truncations', which indicts this repo's own retry-to-16k/64k arms. NOT merged directly because it contests the guide's headline recommendation and is the weakest-sourced item in the set: relayed via screenshot, no published raw data, no handle-confirmed writeup, unlike the lab repo and harness behind the other stacks. Wants review and ideally an independent replication before it changes what the guide tells people to do.
Coherence audit of every card claim against the guide found the card asserting more than main supports: 1. The verdict still said thinking 'helps single-turn codegen', which is the fourth-stack claim I reverted from main pending review in #10. The card was publishing an unreviewed finding as fact. Now states the supported version: net-negative on held-out behavioral work and long agentic loops. 2. The 'Pin + cap' row lost its cap guidance when I reverted the Apollo budget row, leaving a label that promised something the row no longer said, while the guide still (correctly) tells you to set your own ceiling. Restored. 3. '0 runaway loops with pin+cap' was unscoped on the card while the guide carries BlackwellBoy's own scope limit. That matters more now that his gate study reproduces loops 7/10 with the gate open. Card now says zero UNBOUNDED-GENERATION loops, while the thinking gate stayed shut. Also verified the card does not leak any of the five reverted Apollo claims. 15 shared claims now match the guide.
|
Taking the two questions you flagged for challenge, in the order I can actually help with. 1. Task shape: our grid varied shape and apparatus separately, and it supports your inferenceYou wrote that an all-coding benchmark never varied shape, so it measures apparatus rather than shape, and asked for that to be confirmed or shot down. Our C0 to C9 grid crossed four task shapes against ten apparatus levels, 40 samples per condition, so it can separate them. Shape is real and it is large. Firing rate across the whole grid, all conditions pooled:
Summarization never fired once in any condition, including with no system prompt at all. That is a floor no amount of apparatus explains. But code specifically is a conjunction, not a shape effect:
Code fires fully with no apparatus, collapses to zero under a bare named persona, and comes back to 10/10 once the full agent prompt is present. So "coding-shaped tasks suppress" and "coding-shaped tasks reason most" can both be true readings of the same model depending on which single apparatus level you happened to hold fixed. Which is your inference, confirmed: a benchmark that holds shape constant at coding and carries no apparatus is sitting at our C0 cell, where code fires 10/10. It is measuring the apparatus axis at one point, not the shape axis. It cannot see the C4 collapse and it cannot see the summarization floor. Raw: 2. On the sampling confound in the +2.64You already caught that thinking-on ran t0.7 and thinking-off t0.6. I would add that the direction of that confound is not obviously benign for the claim. On a verifiable-answer codegen benchmark the lower-temperature arm is usually the one you would expect to score higher, so a lower-temperature arm losing by 2.64 points is not the shape you would predict from temperature alone. That argues the effect is probably real rather than manufactured, but it also means the magnitude is not clean, and +2.64 is small enough that a two-variable comparison cannot really defend it as a number. Scope-limited merge with the magnitude marked provisional seems right to me, if replication does not arrive. 3. On cap-hits as failures rather than truncationsThis one I can support from a different angle, and it generalizes past Laguna. We ran a six-requirement acceptance-criteria coding task on Qwen 3.6 35B-A3B at a fixed 4096 ceiling, thirty times across three prompt conditions. It returned completely empty content, having spent the entire budget inside the reasoning block, in 28 of 30 runs. Laguna on the same task and ceiling did it 9 times in 30. Not wrong answers, no answer at all. That is the same failure your §5f describes, on a different model and a different serving stack, which I think promotes "treat cap-hits as failures" from a Laguna harness note to a general benchmarking rule. It also matches the degeneration-loop reading rather than the needs-more-budget reading: 4096 is not a stingy ceiling for that task, and more of it was not going to help. Raw: 4. On the regime split being real versus being the benchmarkI do not think our data settles this and I would not claim it does. What I can say is that the two regimes differ on more than task shape: single-turn codegen with a verifiable answer has no apparatus and no accumulated context, and your behavioral and long-agentic arms have both. Since apparatus is now an established axis with a large effect, "regime" and "apparatus dose" are currently confounded with each other in the comparison. Splitting them would need the codegen benchmark re-run with a full agent prompt attached, which is a different experiment from the one anybody has run. Worth stating in the section if it lands, because otherwise the regime table reads as a property of task type when it may partly be a property of how much scaffolding each regime happens to carry. 5. Two small thingsNothing to do with the claim, just noticed while reading the diff.
On replicationIt is running now, on full-precision Laguna S 2.1 NVFP4 at rev
Results to this thread when they land, expected within a day, either way. If it comes back with ON ahead on our stack too, that is the replication you asked for. If it comes back flat or reversed once temperature is controlled, that is worth knowing before this changes Quickstart #1, and I will say so just as plainly. Raw data published the way everything else in our lab repo is. |
|
Correcting myself on one of the two small things above, before you spend time on it. I said the card's footer still says three independent stacks while the row above says four. That is wrong, and it is my misreading rather than a bug in your diff. The two numbers count different sets: the header counts every stack including your own, the footer counts the ones independent of you. Main is "three stacks" in the header against "2 independent stacks" in the footer, and this PR moves both together to four and 3. Same convention, correctly updated. Nothing to fix. The second bullet still stands as far as I can tell: a 12k ceiling makes truncated-think turns more likely than 16k did, so Quickstart row 4 and the §5g failure interact, and a pointer between them would be worth one clause. Apologies for the noise on the first one. |
|
will fix the two small things thanks @Blackwellboy |
Correcting my own over-correction from the previous commit, which stood for one commit. I wrote there that my "an all-coding benchmark with no apparatus measures apparatus rather than shape" inference was WRONG. @Blackwellboy's reply on #10 confirms it: his grid crossed four task shapes against all ten apparatus levels, so it can separate the axes, and a shape-constant no-apparatus benchmark sits at the C0 cell where code fires 10/10. It cannot see the C4 collapse or the summarization floor. Inference restored, with the brief retraction noted rather than hidden. The resolution has two parts and I had only written the second: Shape is a large independent effect. Pooled across all ten conditions, 100 samples per shape: math 92/100, code 62/100, reasoning 47/100, summarization 0/100. Summarization never fired once under any condition including no system prompt, which is a floor no apparatus explains and makes the task itself the strongest single suppressor in the study. Code specifically is a persona conjunction: 10/10 bare, 0/10 under a bare named persona, 10/10 again under a full agent prompt. That is why two stacks could disagree, each held a different single apparatus level fixed. All numbers re-derived from grid_turns.jsonl. Also adds a cross-model pattern: an empty response at a token cap is a failure, not a truncation. Same task and ceiling, Qwen3.6-35B-A3B returned empty content 28/30 and Laguna 9/30, verified per-sample from his published criteria logs. Carries two harness rules (score cap-hits as failures, log the rate per arm) since dropping them inflates whichever arm degenerates most, normally the thinking-on arm. Card "Task gates it" row now carries the pooled numbers; PNG regenerated.
…ller Per @Blackwellboy's review point on this PR. The 12k recommendation and the truncated-think failure pull in opposite directions and I had them sitting in the same row with no pointer between them. 12k comes from the p95 of SINGLE-TURN outputs. A tighter cap makes a turn more likely to exhaust its budget inside the reasoning block, which is the §5f wire failure. In a benchmark that costs one scored sample. In an agent loop the server rejects the malformed assistant message and every following turn fails identically, silently, with nothing to show for it. So the advice is now split: ~12k and score cap-hits as failures for single-turn work; keep headroom for agent loops AND make the loop survive a capped turn rather than relying on the cap never being hit. Noted that the kwarg removes the failure at the source (0/15 vs 2/15 cap-hits). Quickstart row 4 and the card budget row carry the same split.
|
Thanks, this is the most useful review comment the repo has had. Taking your points in order, then what I am running. 1. Task shape: inference confirmed, and I had just talked myself out of itYour pooled table is the thing I did not have. I re-derived it from
Awkward timing on my end and worth stating plainly: about an hour before your comment landed I shipped a commit that declared this inference wrong and retired it, on the strength of your C0-vs-C4 code split alone. That was an over-correction, it stood for exactly one commit, and it is fixed in #12 with the retraction visible rather than quietly reverted. The two parts do not compete: shape is a large independent effect (the summarization floor is a floor no apparatus explains) and code specifically is a persona conjunction. A shape-constant no-apparatus benchmark sits at C0 where code fires 10/10, so it cannot see the C4 collapse or the summarization floor. That is your confirmation, folded. Everything from this comment that touches §2 is in #12, not here, since #12 is the section rewrite. Your review on that one is welcome, particularly on whether "task and persona interact" is more weight than one grid should carry. 2. The +2.64 and the direction of the confoundYour framing is better than mine and I have adopted it: on a verifiable-answer codegen benchmark the lower-temperature arm is normally the one you would expect to score higher, so a lower-temperature arm losing by 2.64 is not the shape temperature alone predicts. That argues the effect is real while the magnitude stays unclean. Agreed on scope-limited merge with the magnitude marked provisional if replication does not land. 3. Cap-hits as failures: promoted to a cross-model patternI verified both lanes per-sample from your published logs rather than from your table: 4. Regime and apparatus are confoundedThis is the sharpest point in your comment and I had not seen it. Single-turn codegen with a verifiable answer carries no apparatus and no accumulated context; our behavioral and long-agentic arms carry both. Since apparatus is now an established large-effect axis, "regime" and "apparatus dose" are confounded with each other in the very comparison this PR is about, so the regime table may partly be measuring scaffolding rather than task type. It will be stated in the section if this lands. Separating them needs the codegen benchmark re-run with a full agent prompt attached, which nobody has run, and I am noting it as the follow-up experiment rather than pretending this PR settles it. 5. The two small thingsNoted on the footer, no apologies needed, and your own correction is right about the convention. Unrelated: main's footer now reads "2 independent replications" instead of "2 independent stacks", same convention, just less ambiguous. The second bullet is fixed, in b199529 on this branch. You had the section letter off by one (§5f is truncated-think, §5g is the gfx1151 serving config) but the interaction you identified is real and I had the two recommendations sitting in the same row with no pointer between them. The advice is now split by regime, because 12k is a p95 of single-turn outputs:
What I am running now, on a third stackSince you are covering NVFP4/vLLM, I have put ours on the other side of the serving axis so the two runs are not the same experiment twice: Q4_K_M GGUF on poolside's own llama.cpp fork at Controls, matching what your review and mine both asked for:
Then a follow-up pass re-running only the cap-hitters at a raised ceiling, which is the direct test of "raising the budget does not convert cap-hitters into passes." If they convert, that claim is wrong on this stack and I will say so. Results and raw JSONL to this thread either way, including if it comes back flat or reversed. Two independent replications on two different quants and two different serving stacks is a much better basis for changing Quickstart #1 than one screenshot was. Note for whenever this does merge: this branch predates today's §2 rewrite and the patterns additions, so it needs a rebase on main first. Its card still carries the retracted dose-curve row. |
|
Glad the pooled table was the missing piece, and respect for leaving the over-correction visible instead of quietly reverting it. That is the standard this repo keeps setting. Your llama.cpp run design looks right to me, and the deliberate no-system-message call is exactly the control that matters now. Coordination point so our two runs stay maximally comparable: mine is also no system message (deliberately the C0 cell), temperature fixed identically across arms, 12,288 ceiling, arms interleaved with nonces, all 164 HumanEval+ problems, evalplus scoring, per-sample cap-hit and extractable-answer logging. So the only intended differences between our runs are quant and serving stack, NVFP4/vLLM against Q4_K_M/llama.cpp, which is what makes the pair informative. If yours and mine land the same sign on opposite stacks, that is about as settled as this question gets short of Poolside's own data. Results to this thread as soon as my run completes, raw JSONL alongside as usual. Two small notes on the cap-hit pattern entry, since I re-derived it from my raw before saying anything. The numbers are all clean: 28/30 and 9/30 both reproduce, both really are three prompt conditions per model, and every one of those empty responses really did stop at exactly 4,096 completion tokens. Nothing to fix there. One sentence overreaches slightly though. The entry says "the same request under a lower apparatus dose completes fine, which is what separates a loop from a genuine truncation." That holds for Laguna and not for Qwen. Per condition:
Laguna behaves the way the sentence describes, fine at low dose and collapsing under the agent prompt. Qwen returns nothing at every dose including the bare prompt. If anything that strengthens the entry's conclusion rather than weakening it, because a failure that does not depend on apparatus at all is even harder to read as a truncation, but the stated discriminator is Laguna-only and I would not want it carried as a general rule. Raw: Also minor, you described that entry as being on main. It is currently in #12 rather than main, so it will land when that merges. Flagging only so it does not get lost if #12 gets reshaped. |
Interim from our replication, 77/328 requests inThird stack, other side of the serving axis from @Blackwellboy's NVFP4/vLLM run: Laguna S 2.1 Q4_K_M GGUF on poolside's own llama.cpp fork at No pass-rate numbers in this comment. Scoring is done by the official Paired tasks with both arms complete (n=38):
Four things worth putting on the record now, because two of them bear on the methodology half of this PR rather than on the +2.64. 1. 2. Zero cap-hits at the 12,288 ceiling, so far. Not one request in 77 has hit it. If that holds to 328, the cap-retry pass I built has nothing to re-run, and the honest reading is not "the claim is wrong" but "the ceiling is not the binding constraint on this stack", which is a different and weaker statement than the guide currently makes. 3. Our p95 is roughly a third of theirs, which is the interesting divergence. @apollo-mg reported p95 10,152 and median 1,945, and derived ~12k from it. We are measuring p95 3,093, median 1,069, longest single response 5,025. Same benchmark, same task shape, no system message in either case. So the ~12k figure looks like a property of their stack and sampling rather than of the model, and a guide-level "use ~12k" recommendation sourced from one stack's p95 is doing more extrapolating than the data supports. Worth reconciling against @Blackwellboy's run when it lands, since he is also using 12k. 4. The cost gap is large and has to be part of the verdict. ON is spending 7.7x the output tokens (1,069 vs 138 median) and 3.8x the wall clock (147s vs 39s). Whatever pass-rate delta comes out, that is the price it has to beat, and a +2.6 point gain bought at 7.7x tokens is a different recommendation from a +2.6 point gain that is free. The current PR text does not price it at all, and it should. Incidental cross-check on @Blackwellboy's grid, free from this run: 164 bare code-shaped prompts with no system message and thinking on is the same cell as his C0 code, which fired 10/10. We are at 34/38 (89%) and the four that did not fire are Timeline: 389 request-minutes of work left at 4 concurrent slots, so ~1.6 hours to all 328 rows, then official evalplus scoring on both arms plus the cap-retry pass if anything hit the ceiling. Numbers and raw JSONL to this thread after that, including if it comes back flat or reversed. I am deliberately not raising parallelism to speed this up: the model is 71 GB resident on a 121 GB unified-memory box, and doubling the KV allocation to get more slots is how that box starts swapping and stops being a valid measurement. |
) (#12) * laguna §2: the thinking gate is two axes, and task shape is resolved Fixes the inference @Blackwellboy flagged in #11, and takes his offer to replace the cross-stack composite curve with his single-stack one. Firing probability and reasoning length move independently, sometimes in opposite directions. The decisive case is his C7 vs C8, which differ by exactly one thing: adding tool schemas cut median reasoning 62% (745 to 282) while firing went UP 12 points (24/40 to 29/40), making it the second-highest firing condition in the grid. My line calling tools a "major suppressor, which is why a maximally coding-shaped benchmark with no apparatus reasons the most" was wrong in both halves and is gone. The bare-prompt conclusion survives on the no-apparatus evidence alone, but not for the reason I gave. The curve is now his ten points: one stack, one revision, one prompt set, n=40 each, varying only apparatus. Our rows and Defilan's are demoted to corroborating the ordering, since absolute levels are stack-specific (his bare 75% vs our ~50%). Numbers re-derived from his published grid_turns.jsonl rather than transcribed. Task shape is no longer marked contested. Both earlier stacks reproduce inside one grid because task and persona interact: code fires 10/10 bare, 0/10 under a named professional persona, 10/10 again under a full agent prompt. That retires my "it measures apparatus rather than shape" guess. Also folded: summarization never fired once in 0/105 attempts under any condition, the strongest single suppressor in the study, and math is stickiest at >=9/10 everywhere except the 10-rule block. §3 corrected to match (coding tasks do not suppress on their own) and the persona lever noted as flooring at 3/40, not 0. Grid size corrected to 400 grid turns, 450 logged total. Card and PNG regenerated. * task shape: restore the confirmed inference, add the pooled shape data Correcting my own over-correction from the previous commit, which stood for one commit. I wrote there that my "an all-coding benchmark with no apparatus measures apparatus rather than shape" inference was WRONG. @Blackwellboy's reply on #10 confirms it: his grid crossed four task shapes against all ten apparatus levels, so it can separate the axes, and a shape-constant no-apparatus benchmark sits at the C0 cell where code fires 10/10. It cannot see the C4 collapse or the summarization floor. Inference restored, with the brief retraction noted rather than hidden. The resolution has two parts and I had only written the second: Shape is a large independent effect. Pooled across all ten conditions, 100 samples per shape: math 92/100, code 62/100, reasoning 47/100, summarization 0/100. Summarization never fired once under any condition including no system prompt, which is a floor no apparatus explains and makes the task itself the strongest single suppressor in the study. Code specifically is a persona conjunction: 10/10 bare, 0/10 under a bare named persona, 10/10 again under a full agent prompt. That is why two stacks could disagree, each held a different single apparatus level fixed. All numbers re-derived from grid_turns.jsonl. Also adds a cross-model pattern: an empty response at a token cap is a failure, not a truncation. Same task and ceiling, Qwen3.6-35B-A3B returned empty content 28/30 and Laguna 9/30, verified per-sample from his published criteria logs. Carries two harness rules (score cap-hits as failures, log the rate per arm) since dropping them inflates whichever arm degenerates most, normally the thinking-on arm. Card "Task gates it" row now carries the pooled numbers; PNG regenerated. * task shape: downgrade to prompt-level, per the study author's weight limit @Blackwellboy re-derived the whole diff back from grid_turns.jsonl (43 checks, zero mismatches) and then put the weight limit somewhere I had not looked. It is not sample size and not apparatus coverage. It is that shape is confounded with prompt identity: each task type in that grid is ONE fixed prompt template repeated with a nonce prefix, not 40 different problems. Verified rather than taken on trust: across all 40 condition-by-task cells the within-cell prompt-token spread never exceeds 4 tokens, which is the nonce tokenizing differently and nothing else. So n=100 per shape is 100 repetitions of one prompt and does not buy category-level generality. The summarization floor is the most load-bearing claim and the most exposed. That prompt is also structurally unlike the other three: at C0 it is 288 prompt tokens against 116 to 124, and it is the only one supplying a passage to condense. So "summarization never fires" and "a prompt handing the model a long passage to condense never fires" are not separated by this grid. His lean, which I share, is the task reading, because the floor survives all ten apparatus levels. Recorded as not settled. Section heading softened from "no longer contested" to resolved for the tested prompts, with category-level generalization pending prompt variation. The two-stack reconciliation is unaffected and now says why: it only needs the same prompt behaving differently under different apparatus, which is what the grid shows. Also records the median convention (median_high, matching their published summary.json; an averaging median lands a few tokens lower on C1, C2, C4, C7) so a recomputation does not read as a mismatch. Roadmap gains the cheap follow-up: re-run with several distinct problems per shape to settle the category question. Card carries the caveat inline.
|
Figured I'd jump in with what I've got so far. Laguna-S-2.1 test data — Apollo fleetRaw per-sample data behind the numbers quoted in #10. Hardware (all runs): quad Tesla P100-PCIE-16GB (sm_60), 1063 MHz / 150 W, single node. 1. Thinking ON vs OFF — HumanEval+ full 164, K=3 (492 samples per arm)Each arm at its own card-recommended sampling.
Files: Thinking-ON is inferred, not directly measured, for the t0.7 arm. That run predates the 2. Stopping-rule observationWRONG counts are near-identical between Laguna-ON (30) and a much larger model run on the Laguna produced no extractable code on 11/492 samples; Puzzle on 1/492. Conditioning on samples that produced code, the 3.05-point gap narrows to 1.16 pt Counter-evidence against over-reading it: forcing termination via 3. Persona × tools factorial15 HumanEval+ problems (every 10th), K=1, t0.7 / top_p 0.95 / top_k 20, thinking at Persona string:
Independent composition would predict 1.09 × 0.92 ≈ 1.00×; observed 0.39×. Robustness (n=15, so this matters):
The persona row's ratio of means is 1.51×, which reads as "+51 % reasoning." That is one Files: CONFOUNDS AND SCOPE LIMITS
Files |
|
This is exactly what the thread needed, thanks for shipping the raw. The Q2_K_XL disclosure is the big one: it means the three replications now span the quant axis by accident, yours at 2 bit, Tom's at Q4, mine at NVFP4, all on HumanEval+ both arms. If the sign survives all three, that's a model property. If it fades with quant quality, that's its own finding. Mine also logs thinking firing directly per sample (the plumbing yours predates), so we'll get a measured rather than inferred ON arm at the high end. Results shortly, the run is nearly complete. One note from adjacent data: your caution on cap-hitters as degeneration is warranted. On Qwen the identical symptom turned out to be pure truncation, budget converts 8 of 10 empties to valid at 8192 with zero degeneration in any tail. Different model, but it says the symptom has more than one cause and per-model checking is right. |
|
Shared your response with my agent, got a test queued up. Opus 5: HumanEval/47 hit the 16 k ceiling twice and passed both times — extractable, correct code, then it kept generating until the cap. So finish_reason=length isn't a failure signal on its own, and my config-doc line saying cap-hits must be bucketed as failures was wrong. Fixed, with the correction visible rather than reworded away. The full breakdown of the 12 cap-hits: │ cap-hit │ bucket │ problem pass_frac │ 6 of the 10 truncations sit on problems Laguna never solves in any sample — which is consistent with "budget won't help" but doesn't demonstrate it, because a never-solved problem is also never solved with more budget. Both hypotheses predict the same thing there. The 4 truncations on solvable problems — /44, /90, /118×2 — are the only place the hypotheses diverge, and that's the population BlackwellBoy's test has to target. The decisive experiment is now specified and ready: re-run those 8 problems at a much larger budget and watch whether the 4 solvable ones convert. One wrinkle worth knowing — our 16 k ceiling wasn't a policy choice, it's the server's -c 32768, so testing this needs Laguna relaunched with a bigger context, not just a bigger max_tokens. If truncations convert, our stopping-rule recommendation is wrong for this model and the fix is simply more budget. If they cap again at 48 k, degeneration is confirmed. It's queued behind stage 2 on .194, same as the loop detector — and it's the better experiment of the two, since it has guaranteed positives where the detector test didn't. Worth noting what BlackwellBoy set up by accident: the three replications now span 2-bit, Q4_K_M, and NVFP4 on the same benchmark and both arms. If our +2.64 survives at NVFP4 with directly-measured firing, it's a model property; if it fades as quant quality rises, that's a cleaner finding than the original. Either outcome is worth more than our single point, and his run gets the measured ON arm that ours could only infer. |
|
Replication results, as promised. Full-precision Laguna S 2.1 NVFP4, rev Verdict firstThe +2.64 does not replicate once the temperature confound is removed. On HumanEval+ — the metric the claim was made on — thinking ON scored 89.84% vs OFF 90.85% (mean of 3 seeds; ON −1.02). Paired per problem it is simply flat: 10 problems favor ON, 13 favor OFF, 141 tied. On base HumanEval the delta leans the other way (ON +1.22). Both deltas are inside seed noise, and both are smaller than the gap between scoring variants of the same benchmark within one run. So on a second stack, with one variable instead of two, the accuracy half of the single-turn-codegen row is flat-to-reversed, not +2.6. The temperature difference appears to have been the whole accuracy story — consistent with your own review note that a lower-temperature arm losing by 2.64 was not the shape temperature alone would predict. What DOES replicate — both of the methodology claims
Design (the confound eliminations your review asked for)
Numbers
That last column is worth a sentence: on this stack thinking ON costs ~11× wall clock for a flat accuracy result. Whatever the regime table ends up saying, "ON for single-turn codegen" is not free. What I'd suggest for the merge decision
Raw data: per-sample JSONL (984 rows: full content + reasoning, cap-hit, extractable, compression ratios, seeds, wall times), driver, and analysis script — One serving note for §2, found while running this: the stray leading |
|
Respect for correcting that in public with the mistake visible, that's the standard this thread keeps setting. Your /47 case is a good catch and it's exactly why my run bucketed on zero extractable code rather than on finish_reason. A cap-hit that still produced correct code is not a failure and should not be scored as one. I had the same case and it went the same way: HumanEval/32 seed 2 hit my 12,288 ceiling with extractable code, tail compression 2.88 which is the lowest ratio in my whole cap population, and it was not scored as a failure. That is 1 of my 15 cap-hits, and the other 14 returned zero code. The two criteria are genuinely not interchangeable on my data either: 22 of my ON runs had no extractable code while only 15 hit the ceiling, so bucketing on finish_reason would have both over-counted and under-counted different failures. On the discriminating question, my NVFP4 run has partial data for you already. p95 output on the ON arm was 6,763 tokens against a 12,288 ceiling, so most of that arm was nowhere near budget limited, and 14 of my 15 cap-hits returned zero code with the entire budget consumed inside the think block, tail compression 2.88 to 142.86. The looping concentrated on specific problems rather than spreading: HumanEval/116 and HumanEval/132 capped in all three seeds, accounting for 6 of my 15 cap-hits across 10 distinct problems. That reads as degeneration rather than starvation at this quant. The more useful part for your queued experiment: your three discriminating problems already have NVFP4 data.
So two of the three never approach the ceiling at NVFP4 and the third caps once in three. I would not read that as settling your question, because quant and ceiling both differ between us, but it is consistent with quant quality rather than budget being what moves those specific problems. Worth noting the overlap the other way too: five problems cap on both of our stacks despite the 2-bit versus NVFP4 gap, /76, /116, /118, /132 and /145. The same problems degenerating at both ends of the quant axis is a stronger argument for problem-specific degeneration than either of us could make alone. Your four solvable-problem truncations are still the cleanest test, and the 48k re-run will settle it better than anything I ran. Good catch on the ceiling really being the server's Full numbers and the 984-row raw are in my thread comment above and in the repo if useful for comparison. |
|
I'll probably need to grab another quant or two. While our test is running, here's my agent's assessment: Opus 5: Two things worth having, one of which corrects my correction. Our harness was already right; my prose was the error. Checked directly: finish=length fires 12 times, src=none (no extractable code) fires 11, and our TRUNCATED bucket requires src=='none' — which is exactly the criterion he's advocating. That's why /47 scored PASS. The bug was in what I wrote in the config doc, not in what the harness did. And his NVFP4 numbers plus ours decompose the cap population in a way neither set does alone: │ problem │ our Q2 (16 k ceiling) │ his NVFP4 (12,288 ceiling) │ reading │ │ /44 │ cap 1/3, ext 2/3, pf 0.67 │ 0 cap-hits, ext 3/3 │ quant-driven │ │ /90 │ cap 1/3, ext 2/3, pf 0.67 │ 0 cap-hits, ext 3/3 │ quant-driven │ │ /118 │ cap 2/3, ext 1/3, pf 0.33 │ 1 of 3, ext 2/3 │ intermediate │ │ /116 │ cap 2/3, ext 1/3, pf 0.00 │ caps all 3 seeds │ problem-intrinsic │ │ /132 │ cap 1/3, ext 2/3, pf 0.00 │ caps all 3 seeds │ problem-intrinsic │ The ceiling difference makes this stronger, not weaker. His ceiling is lower — 12,288 vs our 16,000 — so all else equal he should cap more often. That /44 and /90 cap zero times at NVFP4 despite the tighter budget is real evidence those two are fixed by quant quality. Meanwhile /116 and /132 cap harder at full precision than at 2-bit (3/3 vs 2/3 and 3/3 vs 1/3). Those aren't quantization damage — they're something about the problems themselves that survives a total change of stack, engine, and precision. So the five-problem overlap he identified isn't one phenomenon. It's at least two: quant-sensitive degeneration that better weights fix, and problem-intrinsic degeneration that nothing so far touches. That also sharpens our 48 k test — it's the only clean budget-only manipulation at fixed quant either of us can run, and it should be read per-problem rather than pooled, because those two classes will answer differently. One discrepancy worth flagging before the "same order of magnitude" framing sets: his no-extractable-answer rate is 22/492 against our 11/492, but the composition differs sharply — his is 14 cap + 8 finish=stop with no fenced code, ours is 10 cap + 1. An 8× difference in the stop-without-code population is more likely an extraction-criterion difference (evalplus sanitize vs our extractor) than a model difference. If those rates get compared directly, that needs saying. |
Replication results: third stack, Q4_K_M on poolside's own forkRun complete, 328/328 requests, 0 errors. Scored by the official evalplus 0.3.x scorer, not by anything I wrote. Verdict firstThe +2.64 does not replicate here either. It reverses. On HumanEval+, thinking ON scored 88.4% against OFF 90.9%, so ON is 2.5 points worse, and on base HumanEval it is also behind (94.5 vs 95.1). That is the opposite sign to the claim, on the same benchmark, at a third quantization, with temperature held identical. Combined with @Blackwellboy's NVFP4 result, both runs that removed the temperature confound came back flat-to-negative, at two different temperatures and two different engines:
The one arm that shows the gain is the one where the two arms ran at different temperatures. That is now two independent refutations rather than one, and our run is deliberately the awkward middle of the accidental quant axis @Blackwellboy noticed: same engine as Apollo, same hardware and same 12,288 ceiling as BlackwellBoy. So the divergence is not explained by engine or by ceiling. Curiosity worth recording: all three runs land on 90.85% for one arm or the other. Apollo's ON, BlackwellBoy's OFF, and our OFF are the same number to two decimals. That is 149/164, and it is almost certainly a coincidence of a saturated benchmark rather than anything meaningful, but it is the kind of coincidence worth naming out loud before someone reads significance into it. Paired, per problemPass-rate deltas on a saturated benchmark hide how few problems actually move:
Note that 2 of OFF's 10 wins (/47 and /156) are problems where ON burned its whole budget in the think block and returned no code at all. So a chunk of the accuracy gap is not "ON reasoned to a wrong answer," it is "ON never answered." Cap-hits: the degeneration reading is confirmed, and cleanly4 of 164 ON runs hit the 12,288 ceiling. All 4 returned zero extractable code, entire budget spent inside the think block. OFF: 0 cap-hits, 0 unextractable, 0 fired, all 164. The degeneration signature separates on our data too, using unique-line ratio over the reasoning tail:
Per sample:
And ON's p95 completion is 5,261 tokens against a 12,288 ceiling, so as on @Blackwellboy's stack the cap-hitters are not the tail of the length distribution, they are a separate failure mode. Cross-stack problem overlap is the strongest part of this. Of our 4 loopers, 3 also loop on at least one other stack: /132 loops on all three of us, /145 on ours and Apollo's, /47 on ours and Apollo's. That is problem-specific degeneration surviving a 2-bit to NVFP4 quant spread, which none of us could have established alone. On HumanEval/47 specifically, and this cuts against the counter-example. Apollo's /47 capped twice and passed both times with correct extractable code, which correctly prompted the retraction of "bucket all cap-hits as failures." On our stack /47 capped and returned zero code, at the lowest unique-line ratio in our whole population (0.185), and it is one of OFF's 10 paired wins. So /47 is not reliably benign; it is benign at 2-bit/16k and degenerate at Q4_K_M/12k. The right conclusion stays the one both of you converged on independently: bucket on zero-extractable-code, never on The ceilingOur p95 is 5,261 against Apollo's 10,152 and @Blackwellboy's 6,763. So ~12k is comfortable here and nothing non-looping came near it. I want to correct my own interim comment on this thread: at n=38 I reported p95 3,093 and said ours was "roughly a third of theirs." At the full 164 it is 5,261, about half. The direction held but I quoted a number off a partial run and it moved, which is exactly the mistake this thread keeps catching in other people's work. Cost, which the PR still does not price
@Blackwellboy measured ~11x wall clock, we measure 7.5x. Thinking ON costs roughly an order of magnitude for an accuracy result that is flat at best and negative on two of three stacks. Thinking control, third independent confirmation
The 8 that did not fire: /3, /8, /9, /15, /48, /58, /85, /112. Relevant to the merged §2 work: this is the C0 bare-code cell, where @Blackwellboy measured code firing 10/10 and, in this replication, 492/492. We get 95%. Same cell, same absent system prompt, different serving path, consistently a bit lower, which matches the bare-prompt gap already recorded in the guide. Design
Limitations, stated plainly
Still runningThe cap-retry: the 4 cap-hitters re-issued at 32,768, everything else identical, which is the direct test of "more budget does not convert cap-hitters into passes." Per @apollo-mg's warning, our 12,288 ceiling was also effectively a server limit ( Raw JSONL (328 rows: firing, cap-hit, extractability, unique-line ratios, token counts, latencies, full solutions), runner, scorer, and cap-retry script available if useful for comparison. What I now think should happen to this PR
That is @Blackwellboy's recommendation and our data independently supports it. |
Follow-up: what the thread as a whole now establishes, and what it does notSeparate comment on purpose. The one above is our data. This one is about how the four sets of results interact, including where other people's comments change the reading of our claims, not just Apollo's. 1. The strongest result in this thread is not about thinking. It is about method.The original claim went through three stages: published as a screenshot with no data → replicated twice under a fixed confound → reversed. Nothing about that is a knock on @apollo-mg, who published the full raw with a confounds section that flagged the Q2_K_XL entanglement himself, before anyone asked. Worth being precise about what actually moved the number: not sample size, not hardware, not the benchmark. One variable, temperature, held constant. Both controlled runs went negative. They used different temperatures (0.7/0.7 and 0.6/0.6), different engines, different quants, different ceilings, and different sample counts. About the only thing they share is that each arm ran at the same temperature as its partner. That is a cleaner result than a positive one would have been. 2. What upgraded our own claimsOur "thinking is net-negative" headline survives, but the reason I gave for it was too narrow. The guide justified it from held-out behavioral work and long agentic loops, and explicitly conceded single-turn codegen as the regime where ON might win. Three stacks now say ON does not win there either on accuracy. So the recommendation was right and the carve-out I was preparing to add was not needed. That is a case where the review process saved the guide from a change rather than forcing one.
@Blackwellboy's orphan- 3. What downgraded our own claimsOur p95 claim moved under me and I have corrected it on the thread. I reported "p95 is a third of theirs" from a partial run; at full n it is about half. Same class of error as quoting a table instead of re-deriving from raw, which is the thing I have been asking others to avoid. Our K=1 is the weakest methodology in the thread. Both other replications are K=3 / 492 samples. We cannot report flakiness at all, and flakiness is the one axis where ON genuinely wins on both stacks that can measure it. So on the part of the regime claim most likely to survive, we have nothing to contribute, and our accuracy number should be read as a sign agreeing with @Blackwellboy's rather than as an independent magnitude. The task-shape resolution merged in #12 also needs the prompt-identity caveat carried here. @Blackwellboy found that shape in his grid is confounded with prompt identity, one fixed prompt per task type. Our HumanEval+ run is the opposite trade: 164 genuinely distinct coding problems, but only one shape and one apparatus level. Neither design alone separates shape from prompt from apparatus. Together they bracket it: his varies apparatus with prompt fixed, ours varies prompt with apparatus fixed. Both point the same way on the C0 code cell (his 10/10 and 492/492, our 95%), which is more than either shows alone. 4. Where the thread still disagrees, honestlyCap-hits: same symptom, more than one cause, and this is now well established. Three positions that all turn out to be right in their own scope:
The synthesis all three of us reached independently is the operational one: bucket on zero-extractable-code, never on The regime-versus-apparatus confound is unresolved and nobody's data touches it. @Blackwellboy's point stands: single-turn codegen has no apparatus and no accumulated context, our behavioral and long-agentic arms have both. Since apparatus is now an established large-effect axis, "regime" and "apparatus dose" are confounded in the exact comparison this PR is about. Every run in this thread, all four, sits at the no-apparatus end. The discriminating experiment is the codegen benchmark re-run with a full agent prompt attached, and it has not been run by anyone. It is the highest-value open item here and I would rather name it than let the regime table imply it is settled. 5. What I am taking away about running this repoFour people, four stacks, and the two claims that survived contact are the two that were mechanically checkable: a template rendering, and a per-sample bucket rule. The two that did not survive were both single-stack inferences from summary statistics: my dose curve and the +2.64. That is not a coincidence, and it suggests the bar for anything entering a guide should be "can someone re-derive this from published raw on their own hardware," not "is the number large." Both of the guide's biggest corrections today came from someone re-deriving my numbers from raw and finding the reasoning wrong while the arithmetic was right. Worth saying plainly: this thread has been more useful to offlabel than my own testing was. Cap-retry results to follow, and I will report them the same way whether or not they support the reading I have argued for above. |
Cap-retry results: neither hypothesis was right, and the reason mattersPromised follow-up. The 4 cap-hitters from our Q4_K_M run, re-issued at 32,768 (2.7x the original ceiling), everything else identical. Per @apollo-mg's warning this needed the server relaunched at
Read the middle column before the verdict column. Three of the four "converted", but they finished at 2,153, 3,767 and 6,985 tokens, all well below the original 12,288 ceiling they supposedly needed more budget than. More budget cannot be what fixed them. They simply did not loop on the new draw. So the honest finding is a third thing neither side proposed: on this stack the loops are stochastic per sample, not deterministic per problem. Same problem, same prompt, same temperature, same everything except the draw: loops one time, finishes in 2k tokens the next. That reframes both readings:
This is also the honest limitation of my own retry design, and I should have seen it before running it. Re-issuing at temp 0.6 with a fresh nonce draws a new sample; it does not resume the old one. So "converted" conflates "more budget helped" with "did not loop this time." The only reason the result is still interpretable is the token counts, which happen to discriminate: finishing under the old ceiling rules out budget as the cause. A cleaner design would pin the seed per arm, which is what @Blackwellboy's run does and mine does not. The unique-line ratio separates cleanly, which is good news for the loop-detector idea:
Non-loopers land at 0.5 to 0.6 in both the original and retry populations; the persistent looper is 6x lower. That is a usable online signal, and it is the same discrimination @apollo-mg's compression-ratio stratification and @Blackwellboy's 2.9 to 143x tail ratios found independently. Three different metrics, three stacks, same separation. What loop recovery is actually worth, scoredRe-scoring the ON arm with the 4 retried samples substituted in, via the official evalplus scorer:
So a perfect loop-recovery rule is worth +1.2 points on HumanEval+ here, and ON still ends up behind OFF (89.6 vs 90.9). That is the same conclusion @Blackwellboy reached on NVFP4, in his words "it brings ON to parity on HumanEval+, not ahead," now reproduced at a second quant with a different mechanism (resampling rather than detection). Loop recovery is worth building, and it does not rescue the regime claim. Practical consequence for the stopping ruleIf loops are stochastic rather than problem-inherent, then retry-on-loop-signature is a better rule than raise-the-budget, and it is much cheaper: our three recoveries cost 2k to 7k tokens each, against the 32,768 the budget-raising approach spends before giving up on /145. A detector firing on a low unique-line ratio and re-drawing at the same budget would have recovered 3 of 4 at a fraction of the cost. Worth noting the cross-stack texture on /47 one more time, since it is now a three-way: benign at Apollo's 2-bit (capped twice, passed both), degenerate on our first Q4 draw (zero code, ratio 0.185), benign on our second Q4 draw (3,767 tokens, ratio 0.600). Same problem, three characters. That is what stochastic looping looks like, and it explains the apparent disagreement without either party being wrong. Raw: |
This contests the guide's headline recommendation, so it is deliberately not merged. Requesting review and, ideally, an independent replication first. @Blackwellboy @Defilan, you both have stacks that could test this.
The claim
@apollo-mg measured thinking on HumanEval+, 492 samples, and got the opposite sign to our behavioral battery:
falseeliminates truncated-think (0/15 cap-hits vs 2/15)If it holds, this guide has been costing readers about 2.6 points and double the flakiness on single-turn codegen by recommending thinking off across the board, and we never measured that price. That is why I would rather over-review it than ship it fast.
Two methodology findings come with it, and one of them indicts this repo:
Why this is not merged
It is the weakest-sourced item in the current set. It reached me relayed as a screenshot: no published raw data, no logs, no harness, and I have not seen a writeup under the author's own name. Compare @Blackwellboy's lab repo and @Defilan's merged harness, both of which let me check the work. I am crediting
apollo-mgon the strength of being a known contributor to a sibling repo, not on the strength of published data, and that distinction should be visible rather than smoothed over.What would get this merged
Any one of these:
max_tokens,enable_thinkingtrue vs false, a verifiable-answer benchmark, and see whether ON beats OFF. The mergedscripts/thinking-probes/thinking_ab.pyalready does the interleaving, nonce, and control-cell handling.Specific things I would like challenged
check_guides.pypasses. Main currently reverts to the reviewed wording with a pointer here.