a transformer block in ternary weights and whole numbers
Harriett Little. Published 2026-10-06. The program in this page runs in your browser on cint_ref, the reference. Edit it and run it again.
Ternary language models keep each weight as one of -1, 0 or +1, so the matrix products that make up most of inference become additions and subtractions, and the published work on them reports that this costs little in quality at scale. What those models keep in floating point is everything around the products: the scale that brings activations back into range, the layer norm, the softmax and the residual stream. Those parts are where two machines disagree, because a floating-point sum depends on its order and a floating-point function depends on its library. This page writes one complete block with none of that. The weights are bit planes, the scales are shifts, the softmax is base two with a whole exponent, the layer norm uses an integer square root, and the same program produces the same sixteen logits from the reference and from the compiler, to the byte.
what a ternary product is
A row of ternary weights over sixteen inputs is two 16-bit masks, one for the inputs that count positively and one for those that count negatively, with no bit set in both. The product of the row with an activation vector is the sum of the activations under the first mask minus the sum under the second. No multiplication happens, and because the activations are whole numbers the sum is exact and its order is irrelevant. The block has six such matrices, 2,048 weights in all, generated so that every run sees the same ones, and 713 of them are zero.
The first of the six matrices, the query projection, as the program generates it: 85 weights of +1, 91 of 0 and 80 of -1. Row 0, outlined, is stored as the two masks on the right, and its product with an activation vector is one sum less another. Drawn by replaying the program's generator; the replay's count of all 2,048 weights and the 713 zeros among them is checked against the program's first line.data
row
weights, input 0 to 15
0
- 0 0 - - - + 0 0 0 0 + - - + +
1
- + - 0 0 - 0 0 + 0 - + 0 + 0 +
2
0 0 0 + - + + 0 - + - 0 - - - 0
3
- + 0 + 0 0 - 0 + + - + - + + 0
4
0 + + + 0 - 0 + + 0 + - - 0 0 +
5
- + 0 + 0 - - 0 - 0 0 0 0 + + 0
6
+ + + + 0 + 0 + - 0 0 0 0 - + -
7
+ 0 - 0 + 0 0 0 - + 0 + + 0 0 +
8
- + 0 + + + + 0 + + - + - + + +
9
0 - - 0 + - + - + 0 - + 0 0 0 0
10
0 0 - 0 + + - 0 + - + - 0 - - +
11
+ - - - - 0 0 0 + 0 + 0 0 - + -
12
- - + 0 + + - 0 - - - - 0 - + -
13
+ 0 - - 0 0 0 - + + - - - 0 - +
14
- + 0 0 - 0 - + - 0 - - 0 - 0 -
15
+ 0 0 + + - 0 0 + + - 0 - - + +
the parts that are usually floating point
Scaling. After a product the values can be sixteen times larger than a byte holds, so they are brought back by a shift, meaning a division by a power of two and not the >> operator, which rounds down: the program finds the smallest shift that brings the largest magnitude to at most 127 and divides every value by that power of two with a named rounding. This is the role a per-token absmax scale plays in the published models, and it is exact.
Softmax. Attention scores are divided by the square root of the width, which for sixteen is four, and the weights are two to the power of how far each score falls below the best, counted in steps of 4,096, with the exponent rounded to a whole number. The weights are then exact powers of two, their sum is exact, and the weighted mean of the values is one division with a named rounding. The whole exponent is a choice, stated here, not an approximation hidden in a table.
The weight a score gets, relative to the best score, against how far below the best it falls. The program halves it once per 4,096 score units with the exponent rounded to a whole number, so every weight is a power of two; the dashed line is the same base without the rounding. The dots are the ten scores the program computes for its four tokens: all lie within 992 of the best, under half a step, so every weight is 1. Drawn from the program run in cint_ref with one print line added after the weight.data
token
score
below the best
exponent
weight, of 65,536
0
0
0
0
65,536
1
0
0
0
65,536
1
1
233
0
65,536
2
0
143
0
65,536
2
1
0
0
65,536
2
2
705
0
65,536
3
0
759
0
65,536
3
1
992
0
65,536
3
2
0
0
65,536
3
3
992
0
65,536
Layer norm. The mean is a rounded division, the variance is a sum of squares, the standard deviation is the integer square root of the mean square, and each value is centered, scaled to a byte and divided by that root, all with named roundings. The residual stream is plain addition of whole numbers.
The nearest published work is I-BERT, which runs BERT-family models in integer-only arithmetic: 8-bit weights and activations, polynomial approximations for the exponential in softmax and for GELU, an integer square root found by iteration for layer norm, and rescaling by an integer multiply and a shift. Its exponential splits the input the way the softmax here does, into a whole power of two and a remainder, and approximates the remainder with a second-order polynomial; the block here drops the remainder. The aim there is to stay close enough to a trained floating-point model to keep its accuracy, and to run on integer-only hardware. The aim here is smaller: every step is an exact operation with a named rounding, and checked, so the block gives the same bytes on every machine and stops with a record instead of wrapping if a value does not fit. A block built on I-BERT's polynomials would be a different program with the same property, provided its intermediates were checked too.
the program
Width 16, one attention head, a hidden width of 32, a vocabulary of 16, four tokens. Weights are generated, tokens are embedded as small whole numbers, then projections, attention over earlier tokens, layer norm, the two-layer mlp with a rectifier, and the output row for each vocabulary entry. 186 lines of cint, and every line of it is in the box.
block.ciRun the first Run loads Python, about 12 MB, once
via cint_ref (reference written in .py)
stdout
outcome
what the output shows
Sixteen logits for the last token, the index of the largest, and a checksum over all of them. The numbers themselves mean nothing, because the weights are generated rather than trained; what matters is that they are the same numbers everywhere. Change a token and the logits change. Change the order in which the tokens' projections are computed, or the order of the sums inside a product, and they do not. Printed, the attention weights are all equal: every score falls within 992 of the best, under half of one 4,096 step, so each exponent rounds to zero and the attention in this run is the plain mean of the tokens so far; a smaller step, or weights that set the scores further apart, would give weights below one. Change SCALE from 4_096 to 256 and run again: the last token now gives the third a weight of one, the first an eighth, and the second and itself a sixteenth each, and the next token changes from 11 to 13. The test block checks that the shift rule brings every magnitude up to a hundred thousand within a byte.
same bytes from the reference and the compiler
The compiled program writes the same three lines as the reference, byte for byte, and the test passes under both. The one emitted C file was built with four compilers on three machines, each at -O2 (/O2 for MSVC), and run once on each; every build exited with status 0, as a program that completes does:
The box below is the C that cintc emitted for the program as shipped, 3,170 lines for 186 lines of cint, most of it the checked arithmetic written out.
cintc (the cint compiler)
what is not claimed
This is a block, not a model. It is width sixteen, one head, one layer, with generated weights, and it says nothing about the quality of ternary models, which is the published work's result and not this page's. It does not run on a GPU or a TPU; it is plain sequential cint, built through C for CPUs, and the compiler's CUDA path, which runs kernels on an NVIDIA GPU, is not yet a supported backend. The machines in the identity claim are the four in the compiler's receipts. A trained ternary model exported as bit planes would run through the same code with larger constants, and that is the next step, not this one.
limits
The whole-exponent softmax and the byte-range shifts are the program's own quantization choices, and a trained model would choose its own; what carries over is that each choice is an exact integer operation with a named rounding, so it is the same on every machine. Accumulators are I64 and the widest product in the block is about three million, far inside range; the bounds line of a larger model would be written the way the particle paper writes it. The weights fit one 64-bit word per row because the widths are small; wider rows are more words and the same loop.
references
H. Wang, S. Ma, L. Dong and others. BitNet: Scaling 1-bit Transformers for Large Language Models. 2023.
S. Ma, H. Wang, L. Ma and others. The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits. 2024.
A. Vaswani and others. Attention Is All You Need. Advances in Neural Information Processing Systems, 2017.
J. L. Ba, J. R. Kiros and G. E. Hinton. Layer Normalization. 2016.
S. Kim, A. Gholami, Z. Yao, M. W. Mahoney and K. Keutzer. I-BERT: Integer-only BERT Quantization. International Conference on Machine Learning, 2021.