For the first several thousand tokens, every run of the model agreed about what the next token was going to be regardless of backend. Then in later portions of the prompt, backends began disagreeing. Triton was selected as the baseline to simplify upcoming quantization chicanery.
Each 8k-token window contains 250 sampled positions, one probe every 32 tokens. The percentage is the fraction of those probes where the other backends highest-scoring token differed from Triton’s.
Random noise was accounted for by running the same test with the same attention backend multiple times. The logits across runs at every hidden state were bit for bit identical. Meaning this particular divergence comes exclusively from the matrix multiplication and addition operations happening during prefill inside trt/fa2/fi.
Disagreements appeared in clusters and varied with prompt content rather than increasing smoothly with context length. This is not evidence of one universal length at which the model “falls apart” but… we will get there soon…
Now that we have a baseline comparison of interesting prompt fuel, lets dive into…
Repeating the same methodology, we took the BF16 weights and BF16 kv cache baseline above running Triton, and ran the next experiment. What happens when you leave the weights and activations alone, and JUST quantize the kv-cache?
triton-bf16-baseline-kv-cache-top1-lines-8k-prompt2-96k3317×1370 336 KB
Ah, divergence. And this leads us to our first dumpster-fire of the evening: a completely reproducible tool calling error.
Enough top-tokens got flipped during tool calls, we let them play out and while BF16 was fine, int8 kv-cache eventually managed to recover, int4 did not!
friends dont let friendz 2 lol1122×1402 534 KB
This time we are leaving all the kv-caches full size at bf16. We are adding some new players to the game however by comparing:
These 4 quants represent a broad picture of weights and activations. A notable piece of information for our mathnasium is the actual CUDA kernel / GEMM (general matrix multiplication) / MMA (matrix multiply accumulate) instructions being run to calculate the logits for each quant are different:
five-way-quant-bakeoff-top1-8k-prompt2-96k2432×1382 387 KB
The next-token flip results shake out fairly predictably. TheDude (W8A16) mops the floor with everybody, beating first party FP8 (W8A8) and Nvidia(FP4-is-a-Lie) release. In fact, out of the 5 options, Nvidia’s release comes in dead last hitting ~50% token flips by the time we reach 88k context.
Both the NVFP4 and AWQ W4A16 failed to properly close their tool calls and botched Cisco command line syntax (the correct command was ‘show arp’, while they executed ‘show run’), while both FP8 and INT8 were able to complete the correct calls.
In future experiments I will try to explore the impact of using different fused GEMMs for the same weights, this is another interesting source of divergence where sometimes you have to trade precision for speed.
I have quite a few more experiments and observations to post, but require a great deal of parallel GPU time to calculate and record every logit sampled across huge context chains on multiple prompts with dozens of different settings.
If you have specific questions, shoot me a DM or poke me on discord I guess.
Last weekend I ran a broad statistical comparison centered on one dataset and 3 specific scenarios: What happens when you change the attention cuda kernel, what happens when you lobotomize KV cache, and what happens when you compare the base model with two 8bit and two 4bit quants.
The output was centered on probability distributions that were severe enough to result in token flips, top-1 change output. Some of these were absolutely fascinating when viewed in depth, so for part 2 of my evil plan to take over^H^H^H^H^H^H drag the local LLM sins out into the open, the methodology is going to shift.
Rather than stay high level and capture 3% of the logits, I am now going to capture 100% of the logits for the most impactful areas of the workstream: during tool calls. Gentlemen, we need to go deeper…
I built a small visualizer for my massive hypercube of test case logit captures. It shows a parallel stream of output tokens from some number (2-5) comparable runtimes. And when they differ? We branch and follow both.
The only requirement across runs is the dictionary be the same (so I’m staying within the qwen3.x model family) but can be anything.
Low level cuda kernel and NCCL path differences, driver differences, vllm container runtimes, different GPUs, different combinations of multiple GPUs, attention runtimes, different caching, different model quantizations, and in fact… even different models. 3.6 vs 3.8 anyone?
When the token flip happens, we do not stop and yank the wrong model back to the teacher. We let it continue. This forked multiverse of token output shows us where it went after the error and how it diverged!
Lets go all the way down to unstructured tensor space, see a real failed tool call, a token flip caused by the difference in attention back end / cuda kernel:
In this example of a single token flip, the model executes a tool call targeting an interface on a Cisco router: GigabitEthernet0/0/1.201
Flash attention 2 gets it wrong. It targets GigabitEthernet0/1/4 instead.
Then, it runs the wrong command AGAIN in two diverging tool calls:
The correct command (trying to find the owner of a mac address) is show mac address table. FA2 tries to show run its way out of the mess that token flip has gotten it into.
In this next example, the LLM tried to configure a description on an interface. The FA2 token flip failed to execute that task at all.
image909×395 33.2 KB
Remember. this is simply changing one configuration option in VLLM to select TRT/FA2/FI. Same gpu, same os/drivers/software/vllm/prompt/cache/weights/activations… And this is REPEATABLE between runs, bit-identical logit captures!
If a simple runtime difference in cuda kernels caused this to happen in production, it could result in a critical network outage. The positioning is so impressively bad I could not have hoped for a better example of why precision measurement and testing matters!
We see token flips in many scenarios. This is comparing Tensor Parallelism vs single GPU:
At TP1 we get an acceptable tool call, at TP2 it fails, at TP4 it succeeds again. WTF?! (This is USUALLY NCCL’s fault when you debug even further and capture the nccl graphs…)
Across the 5-weight quant-off from the weekend:
We have BF16, FP8, INT8, and W4A16 all getting it right. Only NVFP4 fails this tool call >_>
We have a growing pile of test case captures in a variety of prompts. Most of my lab include network automation so the corpus will improve as I identify and scrub additional workstream sessions out of turnstone.
So far I have detailed 100% captures with forking realities with:
I am working to package up some of the testing tools and dataset into a distributable package people can run on their rigs and report results, as well as a vast run using a couple rented hopper/blackwell/etc. GPUs
I took a small subset of results as the larger corpus coalesces and spit out a quick 5090 quant buyers guide: Qwen 3.8 Quant Selection Guide for RTX 5090
Also, for anyone wondering HOW you go about capturing real workflows for re-use in later testing, you just need a model router:
Work continues in the background. Stay tuned for the much larger readout on configs capabilities and costs.
There is some legit mad science going on in here, I love this, thank you again for posting, so much content 
@grok for each post, summarize it in a single paragraph.
ok, it would be a rude joke to not compliment you on the great work done!
“If a simple runtime difference in cuda kernels” … this reminds me of non-determinism in areas where it actually matters by design. The endless hours spend working backwards from the result toward the cause. It would be really ironic for the fundamental computational principles (commutativity, floating point, defined order) to come back to bite the AI neural networks in the ass. Here the recognition of the problem is delayed, because everyone assumes randomness of outputs (and the input varies too). And somewhere there, over the rainbow, sit pure INT calculations taunting us with reproducible builds results.
Killer write up! Appreciate sharing all your work.
I wanted to raise a couple of points regarding KL Divergence that, based on my understanding, are important to call out:
Also, “vocabulary” was mentioned one of the dials, but I think most people would call it the “tokenizer." It’s splitting hairs a bit, but figured it was worth calling out for those who may not know off hand.
ok, it would be a rude joke to not compliment you on the great work done!
“If a simple runtime difference in cuda kernels” … this reminds me of non-determinism in areas where it actually matters by design. The endless hours spend working backwards from the result toward the cause. It would be really ironic for the fundamental computational principles (commutativity, floating point, defined order) to come back to bite the AI neural networks in the ass. Here the recognition of the problem is delayed, because everyone assumes randomness of outputs (and the input varies too). And somewhere there, over the rainbow, sit pure INT calculations taunting us with reproducible builds results.
So I went down that route too. What IF we just represented the math as integer with no rounding. The plan works well from 4/8 bit math. you can cleanly represent those ranges with smaller/sane data types.
4bit x 4bit dot products fit within int16. 8bit x 8bit fit within int32.
but at half precision (int 16) it falls apart. You need progressively larger datatypes (int64) to not clip off the bits and introduce the rounding accumulating error. And doing int64 math billions of times in mma operations is computationally prohibitive.
Speed seems to be why most people accept the error.
Now, there is a tiny added benefit to the bit trimming and randomness in probabilistic computing.
In deterministic computing AxB+C=# every time. Back to my floating point math though, you might actually want 6.999999 or 7.00001. Models are not programmed, they are grown. And that incredible spark of something coming out of that growth is more like an emergent property diffused from gaussian noise. Speed aside, I wonder what you would lose making a truly deterministic model.
Edit: After thinking about it, if i DIDNT use the term “dynamic range” in this response somewhere, I would have a flood of angry rage from reddit and discord. Yes, I know. Im discounting that because while you could add a scale factor to every tensor in INT16 thats effectively re-creating floating point math with a hat and sunglasses. I dont think anybody has an int64 accumulator in GPU hardware so this is all hypothetical, let alone changes to the matmul in cuda kernels etc. I went looking for int16 models and was surprised but everything looked like an experiment or abandoned idea.
Moving to an int64 accumulator means more accumulator registers, accumulator read/write POWER, add-path width, local routing and forwarding bandwidth, output-tile storage x bandwidth… you COULD use int32 as the accumulator, but as stated before that would still be lossy…
Mostly yes, a lot of things were glossed over as I tried to make it both approachable as well as technically useful… That’s why while I have a huge KLD write up aside from the post it was not a central piece of the data.
Our harness does normalize the captured logits. It converts them to float64 and applies log_softmax, then calculates directional (D_{KL}(P_{BF16}|P_{candidate})), reverse KL, and Jensen–Shannon divergence. We retain per-token values and summarize them within individual output ranges rather than presenting one global average.
BF16 is a numerical-fidelity reference, not an oracle or correctness label. A quantized model can absolutely diverge from BF16 and produce a semantically better answer. Repeating the identical deterministic run tests reproducibility, but does not make BF16 correct; correctness requires labelled answers, executable tool-call checks, or semantic grading across varied workloads.
Our Top-1 percentage also is not a percentage difference between logits. It is the percentage of evaluated output positions where the candidate’s argmax token ID differs from BF16’s:
Top-1 disagreement = changed winner positions / evaluated positions. Top-1 and KL describe different things. KL measures movement of the whole distribution, including changes that leave the winner unchanged. Top-1 disagreement is a discontinuous winner test: an extremely small KL change can flip a 50.1/49.9 decision, while a much larger KL change can leave a dominant winner unchanged.
Top-1 is directly relevant to greedy decoding, but a teacher-forced Top-1 flip is still a counterfactual root, not automatically a different complete answer. That is why we additionally branch from selected flip positions and inspect whether the alternate continuation recovers, changes meaning, or produces malformed or incorrect tool calls.
I will be the first to point out using a random 3% token distribution (part 1) to measure divergence is not a correct overall methodology, but i needed a big picture view. That’s why part 2 went straight down the rabbit hole to what-does-a-top1-flip-mean and why would it matter to you in a specific use case.
There is a clear impact, to both end user perception of a model’s output as well as measured correctness in benchmark results. Just wait for what’s coming next…
The initial H200 runs mostly finished last night, and B200 runs finishing this morning sometime.
The mountain-o-tests has grown wildly out of control, well beyond what a hypercube of logits could ever fit into one post. So I am going to start splitting the next segments into moderately entertaining summaries of the results to hopefully explore some of the many… maaany… interesting findings.
(I might edit this post to include a few more charts when i have a chance at lunch so consider this a preliminary release)
With the release of qwen3.8 many people are flocking to fine tunes labeled as HERETIC! UNCENSORED! ABLITERATED!
So, while they might be able to remove some post-training “safety” (i hate that term) what is the overall impact on their ability to actually do work? If you want to chat about the capital of a certain island nation or certain events in 1989, it will probably do just fine. But what if we let it make those tasty tool calls and run it through the battery of forced teacher decodes?
We grabbed 4 “popular” (by likes and top downloads) tunes of Qwen 3.8, all full BF16 sized not quants:
And compared them to our reference BF16 logits captured on SM120 for Qwen/Qwen3.8-27B ( Qwen/Qwen3.8-27B · Hugging Face )
We also ran a limited W4A16 side experiment for fun.
The first two quants appear remarkably functional while the latter two should probably go on the do-not-use list.
qwen38-derivatives-top1-flip-rate-sp04-sp062880×1620 195 KB
SP04 covers 1,535 assistant-output tokens in six natural ranges. SP06 covers 4,339 tokens in seven ranges, including prose, exact CLI/SQL/code, multi-tool calls, recovery actions, and architecture recommendations.