cint

news

Newest first. Dates are US Eastern.

Development Tree

These changes are in the development tree and have not been tagged in a release. The release notes say what each release includes.

GCC 16 Performance Update: Ledger Loop Outperforms Rust

outcome value
machine intel-core-i7-8700K fedora-linux-44
builds gcc-16.2.1 clang-22.1.8 rustc-1.96.1, each -O2
gcc.ledger.vs-checked-c 0.998x
gcc.ledger.vs-rust 0.927x
vs-rust.at-or-below-1.00x 3/4
vs-checked-c.within-1.10x 3/4
timings.valid 8/8
agreement.checks agree 24/24

The two runtime changes projected in the 2026-10-08 performance entry are now timed on a second machine: an Intel Core i7-8700K running Fedora Linux 44, with GCC 16.2.1, Clang 22.1.8 and Rust 1.96.1. After both changes, cint's ledger loop under GCC runs in 0.927x Rust's time (1,058,061 ns against 1,141,463 ns, median) and matches hand-checked C at 0.998x (1,060,602 ns).

The benchmark method matches the earlier runs: cint built with cint build's release settings (-O2, built-in overflow helpers, fuel metering off), hand-checked C built by the same compiler with the same flags and its overflow builtins, and Rust with checked_* arithmetic and overflow checks on at opt-level=2. Each ratio is cint's median time over 110 samples divided by the other program's. All eight timings were valid on the first run, and all 24 agreement checks matched the reference.

cint's median time divided by the other program's on the second machine, after the first runtime change and after both
Loop and Compilervs C, First Changevs C, Both Changesvs Rust, First Changevs Rust, Both Changes
Ledger Loop, GCC 16.2.11.096x0.998x1.021x0.927x
Table Accumulation, GCC 16.2.11.413x1.176x1.000x0.830x
Ledger Loop, Clang 22.1.81.003x1.018x1.032x1.045x
Table Accumulation, Clang 22.1.81.000x1.000x1.000x0.996x

After both changes, cint outperforms Rust in three of the four configurations: the GCC ledger loop (0.927x), GCC table accumulation (0.830x) and Clang table accumulation (0.996x). The Clang ledger loop trails Rust by 4.5% (1.045x).

Against hand-checked C, cint matches or beats it on the GCC ledger loop (0.998x) and Clang table accumulation (1.000x), and stays within 1.10x on the Clang ledger loop (1.018x). The one exception is GCC table accumulation at 1.176x, down from 1.413x after the first change. There, GCC 16 compiles the hand-checked C into markedly faster code (322,533 ns, median) than everything else measured: cint (379,384 ns), Rust (456,851 ns) and the same C built by Clang (455,062 ns).

These measurements complement the 2026-10-08 runs. Its figures for GCC 13.3, Clang 18.1, MSVC and Apple Clang stand as measured, and its two GCC 13.3 projections remain untimed on that machine.

Workbench: Differential Debugging and Execution Replay

outcome value
builds gcc-13 clang-18 msvc-19.44 apple-clang-21, each -O0 and -O2
sessions.matching 3/3 in 8/8 builds
debug.diff agree 3/3
debug.replay ok 9/9
receipts.same-in-every-build 12/12

The interactive workbench gains three diagnostic commands. cint explain shows what the compiler knows about a function without running it, including every operation that can fault, with its source position and fault code. cint debug diff runs a saved case as compiled C and on the interpreter, then reports that the two agree or names the first point where they part. cint debug replay runs a recording again to check that it reproduces, and cint debug locate replays a recording straight to its first fault.

Three scripted sessions were run on GCC 13.3.0, Clang 18.1.3, MSVC 19.44 and Apple Clang 21.0.0, each at -O0 and -O2. In all eight builds, every session matched its transcript, all three diffs agreed, all nine replays reproduced, and all twelve receipts had the same identity. The workbench is 13,455 lines of C. Its GPU commands stay off until cint build --lib can build a GPU library.

Interpreter Engine & State Checkpointing

outcome value
conformance.programs 513
conformance.records 1842927
conformance.disagree 0
executors compiled-c interpreter
builds 24 each
case.saved msvc-19.44 x86-64-windows
case.replayed ok 4/4

cint programs can now run directly on the interpreter, cint-interp, with no C compiler involved. The interpreter executes the same semantic IR that the C backend compiles, so a program means the same thing in both modes. On the conformance suite, 513 programs and 1,842,927 records, the interpreter and the compiled C both matched the reference on every record, in 24 builds each across GCC, Clang, MSVC and Apple Clang, with and without sanitizers. The sanitizers reported nothing.

Checkpointing arrived with it. A program's state can be saved, inspected and restored exactly, across operating systems and processor architectures. A case saved on Windows with MSVC replayed with the same outcome and the same receipt on Linux under GCC and Clang and on an Apple M4 under Apple Clang. An edit that changes only whitespace leaves a checkpoint valid. The interpreter is 6,999 lines of C.

Performance: Matching or Beating Hand-Checked C in Six of Eight Measurements

outcome value
vs-checked-c.at-or-below-1.00x 6/8
vs-checked-c.within-1.10x 7/8
vs-rust.at-or-below-1.00x 2/8
orbit-1e8.vs-double 3.93x to 6.88x
compiler.self-build.vs-before-v0.1.0 0.53x to 1.13x

The benchmarks time two loops from a ledger program: checked addition and subtraction in sequence, and checked accumulation over a 256-entry table, 1,000,000 iterations each. cint builds them the way cint build builds any program (-O2, built-in overflow helpers, fuel metering off). The C versions check overflow by hand with __builtin_*_overflow where the compiler has it and are built by the same compiler with the same flags. The Rust versions use checked_* arithmetic with overflow checks on, at opt-level=2. Each ratio is cint's median time over 110 samples divided by the other program's, so below 1.00x means cint is faster. Every variant was checked against the reference before it was timed.

cint's median time divided by the other program's, per toolchain and loop
Build EnvironmentLedger vs CTable vs CLedger vs RustTable vs Rust
MSVC 19.44, Windows 11, Ryzen 9 5900X0.951x0.789x1.238x1.715x
GCC 13.3.0, Ubuntu 24.04, Ryzen 9 5900X1.179x1.012x1.175x2.032x
Clang 18.1.3, Ubuntu 24.04, Ryzen 9 5900X1.000x0.982x0.917x2.012x
Apple Clang 21.0.0, macOS 27, Apple M40.948x0.999x1.059x1.000x

Against hand-checked C, cint is as fast or faster in six of the eight measurements and within 1.10x in seven. The one exception is the ledger loop under GCC, at 1.179x. A change already in the development tree is projected to bring it to about 1.08x; the projection comes from instruction counts and has not been timed.

Against Rust, cint is faster on the ledger loop under Clang, at 0.917x, and ties on the M4's table loop, at 1.000x. Rust leads in the other six, by as much as 2.03x on the table loop under MSVC, GCC and Clang, where the hand-checked C trails Rust by about 2x as well. On the GCC ledger loop, a second change in the tree brings cint to 19 instructions per iteration, the same as Rust and one fewer than the C. Its speed is projected too and has not been timed.

The orbit paper's program steps an orbit 108 times in checked fixed-point integers. In the paper, built with the compiler behind this site's C and timed in single runs, it was 25 to 44 times slower than the same integrator in double. In the development tree, its median over 11 runs is 3.93x the double build's time on Clang, 4.11x on the M4, 4.36x on GCC and 6.88x on MSVC, inside an allowance of 8x. The double build does not agree with itself: built by GCC with and without fused multiply-add, it ends 93 cm apart, while the integer version reaches the same digest on all four compilers.

The self-hosted build, cintc compiling itself, now takes 0.53x its time from the day before v0.1.0 on GCC, 0.68x on Clang and 0.93x on the M4. On MSVC it takes 1.13x, inside an allowance of 1.25x. Each figure comes from one machine and one toolchain.

NVIDIA PTX Code Generation & CUDA Execution

outcome value
backend gpu-cuda
device nvidia-rtx-3090-ti
builds msvc-19.44 gcc-13 clang-18, blocks 32 and 256
conformance.programs 504
conformance.records 1840878
conformance.disagree 0
conformance.refused 8

Kernels written in cint now run natively on NVIDIA GPUs and match the CPU backend exactly: the same output buffers, reduction values, fuel consumed, dispatch counts and fault records. cintc, itself written in cint, generates the PTX.

On an RTX 3090 Ti, the GPU backend ran the conformance suite 24 times, from MSVC, GCC and Clang builds at block sizes 32 and 256, and matched the reference on all 1,840,878 records it compared. It refuses eight cases that use operators the CUDA backend does not lower: saturating add, subtract and multiply, division and remainder, and the three shifts. Those programs are turned away outright rather than run with different results. From Python, a stencil and a matrix multiply written as cint kernels run on the GPU over NumPy and CuPy arrays and PyTorch tensors, on Windows and under WSL2.

Releases

v0.1.1: Maintenance Release

Released the same day as v0.1.0. The cint command now counts every runtime source file in a receipt's runtime identity; v0.1.0 counted 8 of the 14. The Python bridge is included, so tutorial 04 and the matrix example run. The compiler is unchanged. See the release notes.

v0.1.0: Initial Public Release

outcome value
release v0.1.0
conformance.cases 861
conformance.compared 783
specification.editions SPEC-00P.1 SPEC-01P.1 SPEC-04P.1

The first public release includes cintc, a self-hosting compiler written in cint that emits C17 and rebuilds itself byte for byte; the seed compiler in C that bootstraps it; the runtime; the cint command; the reference implementation, cint_ref; and a conformance suite of 861 cases. cintc agrees with the reference on all 783 cases it is compared on.

The release published the first specification editions, SPEC-00P.1, SPEC-01P.1 and SPEC-04P.1, with a glossary and references. Once published, an edition never changes and keeps its address. See the release notes.