cint

a transformer block in ternary weights and whole numbers

Harriett Little. Published 2026-10-06.
The program in this page runs in your browser on cint_ref, the reference. Edit it and run it again.

Ternary language models keep each weight as one of -1, 0 or +1, so the matrix products that make up most of inference become additions and subtractions, and the published work on them reports that this costs little in quality at scale. What those models keep in floating point is everything around the products: the scale that brings activations back into range, the layer norm, the softmax and the residual stream. Those parts are where two machines disagree, because a floating-point sum depends on its order and a floating-point function depends on its library. This page writes one complete block with none of that. The weights are bit planes, the scales are shifts, the softmax is base two with a whole exponent, the layer norm uses an integer square root, and the same program produces the same sixteen logits from the reference and from the compiler, to the byte.

what a ternary product is

A row of ternary weights over sixteen inputs is two 16-bit masks, one for the inputs that count positively and one for those that count negatively, with no bit set in both. The product of the row with an activation vector is the sum of the activations under the first mask minus the sum under the second. No multiplication happens, and because the activations are whole numbers the sum is exact and its order is irrelevant. The block has six such matrices, 2,048 weights in all, generated so that every run sees the same ones, and 713 of them are zero.

the query matrix as +1, 0 and -1 cells, and its first row as two bit masks the query matrix: row j, input i 0 0 15 15 row 0, input 0: -1 row 0, input 1: 0 row 0, input 2: 0 row 0, input 3: -1 row 0, input 4: -1 row 0, input 5: -1 row 0, input 6: +1 row 0, input 7: 0 row 0, input 8: 0 row 0, input 9: 0 row 0, input 10: 0 row 0, input 11: +1 row 0, input 12: -1 row 0, input 13: -1 row 0, input 14: +1 row 0, input 15: +1 row 1, input 0: -1 row 1, input 1: +1 row 1, input 2: -1 row 1, input 3: 0 row 1, input 4: 0 row 1, input 5: -1 row 1, input 6: 0 row 1, input 7: 0 row 1, input 8: +1 row 1, input 9: 0 row 1, input 10: -1 row 1, input 11: +1 row 1, input 12: 0 row 1, input 13: +1 row 1, input 14: 0 row 1, input 15: +1 row 2, input 0: 0 row 2, input 1: 0 row 2, input 2: 0 row 2, input 3: +1 row 2, input 4: -1 row 2, input 5: +1 row 2, input 6: +1 row 2, input 7: 0 row 2, input 8: -1 row 2, input 9: +1 row 2, input 10: -1 row 2, input 11: 0 row 2, input 12: -1 row 2, input 13: -1 row 2, input 14: -1 row 2, input 15: 0 row 3, input 0: -1 row 3, input 1: +1 row 3, input 2: 0 row 3, input 3: +1 row 3, input 4: 0 row 3, input 5: 0 row 3, input 6: -1 row 3, input 7: 0 row 3, input 8: +1 row 3, input 9: +1 row 3, input 10: -1 row 3, input 11: +1 row 3, input 12: -1 row 3, input 13: +1 row 3, input 14: +1 row 3, input 15: 0 row 4, input 0: 0 row 4, input 1: +1 row 4, input 2: +1 row 4, input 3: +1 row 4, input 4: 0 row 4, input 5: -1 row 4, input 6: 0 row 4, input 7: +1 row 4, input 8: +1 row 4, input 9: 0 row 4, input 10: +1 row 4, input 11: -1 row 4, input 12: -1 row 4, input 13: 0 row 4, input 14: 0 row 4, input 15: +1 row 5, input 0: -1 row 5, input 1: +1 row 5, input 2: 0 row 5, input 3: +1 row 5, input 4: 0 row 5, input 5: -1 row 5, input 6: -1 row 5, input 7: 0 row 5, input 8: -1 row 5, input 9: 0 row 5, input 10: 0 row 5, input 11: 0 row 5, input 12: 0 row 5, input 13: +1 row 5, input 14: +1 row 5, input 15: 0 row 6, input 0: +1 row 6, input 1: +1 row 6, input 2: +1 row 6, input 3: +1 row 6, input 4: 0 row 6, input 5: +1 row 6, input 6: 0 row 6, input 7: +1 row 6, input 8: -1 row 6, input 9: 0 row 6, input 10: 0 row 6, input 11: 0 row 6, input 12: 0 row 6, input 13: -1 row 6, input 14: +1 row 6, input 15: -1 row 7, input 0: +1 row 7, input 1: 0 row 7, input 2: -1 row 7, input 3: 0 row 7, input 4: +1 row 7, input 5: 0 row 7, input 6: 0 row 7, input 7: 0 row 7, input 8: -1 row 7, input 9: +1 row 7, input 10: 0 row 7, input 11: +1 row 7, input 12: +1 row 7, input 13: 0 row 7, input 14: 0 row 7, input 15: +1 row 8, input 0: -1 row 8, input 1: +1 row 8, input 2: 0 row 8, input 3: +1 row 8, input 4: +1 row 8, input 5: +1 row 8, input 6: +1 row 8, input 7: 0 row 8, input 8: +1 row 8, input 9: +1 row 8, input 10: -1 row 8, input 11: +1 row 8, input 12: -1 row 8, input 13: +1 row 8, input 14: +1 row 8, input 15: +1 row 9, input 0: 0 row 9, input 1: -1 row 9, input 2: -1 row 9, input 3: 0 row 9, input 4: +1 row 9, input 5: -1 row 9, input 6: +1 row 9, input 7: -1 row 9, input 8: +1 row 9, input 9: 0 row 9, input 10: -1 row 9, input 11: +1 row 9, input 12: 0 row 9, input 13: 0 row 9, input 14: 0 row 9, input 15: 0 row 10, input 0: 0 row 10, input 1: 0 row 10, input 2: -1 row 10, input 3: 0 row 10, input 4: +1 row 10, input 5: +1 row 10, input 6: -1 row 10, input 7: 0 row 10, input 8: +1 row 10, input 9: -1 row 10, input 10: +1 row 10, input 11: -1 row 10, input 12: 0 row 10, input 13: -1 row 10, input 14: -1 row 10, input 15: +1 row 11, input 0: +1 row 11, input 1: -1 row 11, input 2: -1 row 11, input 3: -1 row 11, input 4: -1 row 11, input 5: 0 row 11, input 6: 0 row 11, input 7: 0 row 11, input 8: +1 row 11, input 9: 0 row 11, input 10: +1 row 11, input 11: 0 row 11, input 12: 0 row 11, input 13: -1 row 11, input 14: +1 row 11, input 15: -1 row 12, input 0: -1 row 12, input 1: -1 row 12, input 2: +1 row 12, input 3: 0 row 12, input 4: +1 row 12, input 5: +1 row 12, input 6: -1 row 12, input 7: 0 row 12, input 8: -1 row 12, input 9: -1 row 12, input 10: -1 row 12, input 11: -1 row 12, input 12: 0 row 12, input 13: -1 row 12, input 14: +1 row 12, input 15: -1 row 13, input 0: +1 row 13, input 1: 0 row 13, input 2: -1 row 13, input 3: -1 row 13, input 4: 0 row 13, input 5: 0 row 13, input 6: 0 row 13, input 7: -1 row 13, input 8: +1 row 13, input 9: +1 row 13, input 10: -1 row 13, input 11: -1 row 13, input 12: -1 row 13, input 13: 0 row 13, input 14: -1 row 13, input 15: +1 row 14, input 0: -1 row 14, input 1: +1 row 14, input 2: 0 row 14, input 3: 0 row 14, input 4: -1 row 14, input 5: 0 row 14, input 6: -1 row 14, input 7: +1 row 14, input 8: -1 row 14, input 9: 0 row 14, input 10: -1 row 14, input 11: -1 row 14, input 12: 0 row 14, input 13: -1 row 14, input 14: 0 row 14, input 15: -1 row 15, input 0: +1 row 15, input 1: 0 row 15, input 2: 0 row 15, input 3: +1 row 15, input 4: +1 row 15, input 5: -1 row 15, input 6: 0 row 15, input 7: 0 row 15, input 8: +1 row 15, input 9: +1 row 15, input 10: -1 row 15, input 11: 0 row 15, input 12: -1 row 15, input 13: -1 row 15, input 14: +1 row 15, input 15: +1 row 0, plus mask input 0: bit clear input 1: bit clear input 2: bit clear input 3: bit clear input 4: bit clear input 5: bit clear input 6: bit set input 7: bit clear input 8: bit clear input 9: bit clear input 10: bit clear input 11: bit set input 12: bit clear input 13: bit clear input 14: bit set input 15: bit set inputs 6, 11, 14, 15 row 0, minus mask input 0: bit set input 1: bit clear input 2: bit clear input 3: bit set input 4: bit set input 5: bit set input 6: bit clear input 7: bit clear input 8: bit clear input 9: bit clear input 10: bit clear input 11: bit clear input 12: bit set input 13: bit set input 14: bit clear input 15: bit clear inputs 0, 3, 4, 5, 12, 13 row 0 times x = the x under the plus mask, added, less the x under the minus mask +1 0 -1
The first of the six matrices, the query projection, as the program generates it: 85 weights of +1, 91 of 0 and 80 of -1. Row 0, outlined, is stored as the two masks on the right, and its product with an activation vector is one sum less another. Drawn by replaying the program's generator; the replay's count of all 2,048 weights and the 713 zeros among them is checked against the program's first line.
data
rowweights, input 0 to 15
0- 0 0 - - - + 0 0 0 0 + - - + +
1- + - 0 0 - 0 0 + 0 - + 0 + 0 +
20 0 0 + - + + 0 - + - 0 - - - 0
3- + 0 + 0 0 - 0 + + - + - + + 0
40 + + + 0 - 0 + + 0 + - - 0 0 +
5- + 0 + 0 - - 0 - 0 0 0 0 + + 0
6+ + + + 0 + 0 + - 0 0 0 0 - + -
7+ 0 - 0 + 0 0 0 - + 0 + + 0 0 +
8- + 0 + + + + 0 + + - + - + + +
90 - - 0 + - + - + 0 - + 0 0 0 0
100 0 - 0 + + - 0 + - + - 0 - - +
11+ - - - - 0 0 0 + 0 + 0 0 - + -
12- - + 0 + + - 0 - - - - 0 - + -
13+ 0 - - 0 0 0 - + + - - - 0 - +
14- + 0 0 - 0 - + - 0 - - 0 - 0 -
15+ 0 0 + + - 0 0 + + - 0 - - + +

the parts that are usually floating point

Scaling. After a product the values can be sixteen times larger than a byte holds, so they are brought back by a shift, meaning a division by a power of two and not the >> operator, which rounds down: the program finds the smallest shift that brings the largest magnitude to at most 127 and divides every value by that power of two with a named rounding. This is the role a per-token absmax scale plays in the published models, and it is exact.

Softmax. Attention scores are divided by the square root of the width, which for sixteen is four, and the weights are two to the power of how far each score falls below the best, counted in steps of 4,096, with the exponent rounded to a whole number. The weights are then exact powers of two, their sum is exact, and the weighted mean of the values is one division with a named rounding. The whole exponent is a choice, stated here, not an approximation hidden in a table.

the attention weight against how far a score falls below the best, and the program's ten scores 1 1/2 1/4 1/8 0 weight, relative to the best score 0 4,096 8,192 12,288 16,384 how far a score falls below the best, in score units token 0, score 0: 0 below the best, weight 1 token 1, score 0: 0 below the best, weight 1 token 1, score 1: 233 below the best, weight 1 token 2, score 0: 143 below the best, weight 1 token 2, score 1: 0 below the best, weight 1 token 2, score 2: 705 below the best, weight 1 token 3, score 0: 759 below the best, weight 1 token 3, score 1: 992 below the best, weight 1 token 3, score 2: 0 below the best, weight 1 token 3, score 3: 992 below the best, weight 1 the program's ten scores, all within 992 of the best whole exponent, the program the same base, not rounded
The weight a score gets, relative to the best score, against how far below the best it falls. The program halves it once per 4,096 score units with the exponent rounded to a whole number, so every weight is a power of two; the dashed line is the same base without the rounding. The dots are the ten scores the program computes for its four tokens: all lie within 992 of the best, under half a step, so every weight is 1. Drawn from the program run in cint_ref with one print line added after the weight.
data
tokenscorebelow the bestexponentweight, of 65,536
000065,536
100065,536
11233065,536
20143065,536
210065,536
22705065,536
30759065,536
31992065,536
320065,536
33992065,536

Layer norm. The mean is a rounded division, the variance is a sum of squares, the standard deviation is the integer square root of the mean square, and each value is centered, scaled to a byte and divided by that root, all with named roundings. The residual stream is plain addition of whole numbers.

The nearest published work is I-BERT, which runs BERT-family models in integer-only arithmetic: 8-bit weights and activations, polynomial approximations for the exponential in softmax and for GELU, an integer square root found by iteration for layer norm, and rescaling by an integer multiply and a shift. Its exponential splits the input the way the softmax here does, into a whole power of two and a remainder, and approximates the remainder with a second-order polynomial; the block here drops the remainder. The aim there is to stay close enough to a trained floating-point model to keep its accuracy, and to run on integer-only hardware. The aim here is smaller: every step is an exact operation with a named rounding, and checked, so the block gives the same bytes on every machine and stops with a record instead of wrapping if a value does not fit. A block built on I-BERT's polynomials would be a different program with the same property, provided its intermediates were checked too.

the program

Width 16, one attention head, a hidden width of 32, a vocabulary of 16, four tokens. Weights are generated, tokens are embedded as small whole numbers, then projections, attention over earlier tokens, layer norm, the two-layer mlp with a rectifier, and the output row for each vocabulary entry. 186 lines of cint, and every line of it is in the box.

block.ci   Run   the first Run loads Python, about 12 MB, once

via cint_ref (reference written in .py)
stdout

    outcome
    

  

what the output shows

Sixteen logits for the last token, the index of the largest, and a checksum over all of them. The numbers themselves mean nothing, because the weights are generated rather than trained; what matters is that they are the same numbers everywhere. Change a token and the logits change. Change the order in which the tokens' projections are computed, or the order of the sums inside a product, and they do not. Printed, the attention weights are all equal: every score falls within 992 of the best, under half of one 4,096 step, so each exponent rounds to zero and the attention in this run is the plain mean of the tokens so far; a smaller step, or weights that set the scores further apart, would give weights below one. Change SCALE from 4_096 to 256 and run again: the last token now gives the third a weight of one, the first an eighth, and the second and itself a sixteenth each, and the next token changes from 11 to 13. The test block checks that the shift rule brings every magnitude up to a hundred thousand within a byte.

same bytes from the reference and the compiler

The compiled program writes the same three lines as the reference, byte for byte, and the test passes under both. The one emitted C file was built with four compilers on three machines, each at -O2 (/O2 for MSVC), and run once on each; every build exited with status 0, as a program that completes does:

runmachinestdout SHA-256
cint_refthis pageff8795b040bf53f7a6fa28dfae71907c620f7bf5489eaea2766a26f7ebf35ecc
clang 18.1.3x86-64 Linuxff8795b040bf53f7a6fa28dfae71907c620f7bf5489eaea2766a26f7ebf35ecc
gcc 13.3.0x86-64 Linuxff8795b040bf53f7a6fa28dfae71907c620f7bf5489eaea2766a26f7ebf35ecc
MSVC 19.44x86-64 Windows 11ff8795b040bf53f7a6fa28dfae71907c620f7bf5489eaea2766a26f7ebf35ecc
Apple Clang 21.0.0arm64 macOS 27ff8795b040bf53f7a6fa28dfae71907c620f7bf5489eaea2766a26f7ebf35ecc

The box below is the C that cintc emitted for the program as shipped, 3,170 lines for 186 lines of cint, most of it the checked arithmetic written out.

cintc (the cint compiler)


what is not claimed

This is a block, not a model. It is width sixteen, one head, one layer, with generated weights, and it says nothing about the quality of ternary models, which is the published work's result and not this page's. It does not run on a GPU or a TPU; it is plain sequential cint, built through C for CPUs, and the compiler's CUDA path, which runs kernels on an NVIDIA GPU, is not yet a supported backend. The machines in the identity claim are the four in the compiler's receipts. A trained ternary model exported as bit planes would run through the same code with larger constants, and that is the next step, not this one.

limits

The whole-exponent softmax and the byte-range shifts are the program's own quantization choices, and a trained model would choose its own; what carries over is that each choice is an exact integer operation with a named rounding, so it is the same on every machine. Accumulators are I64 and the widest product in the block is about three million, far inside range; the bounds line of a larger model would be written the way the particle paper writes it. The weights fit one 64-bit word per row because the widths are small; wider rows are more words and the same loop.

references

cite as

Harriett Little. A transformer block in ternary weights and whole numbers. integerc.dev, 2026. https://integerc.dev/papers/ternary/

Harriett Little
October 2026