I asked another tool to check my numbers. It found my bugs, not its own.
Tezcatl, an open-source C and C++ metrics tool, and what cross-checking it against lizard and gcovr turned up

Tezcatl went public last week. It's a command-line tool that reads a C or C++ project's own build configuration and writes a code metrics baseline: lines of code, cyclomatic complexity, documentation coverage, test coverage, and the include graph with its cycles. MIT licence, source on GitHub, docs at narcoleptic-fox.github.io/tezcatl.
That's the announcement. The part worth writing about is how I decided whether to trust it.
A metrics tool is a claim about someone else's code
If Tezcatl says a function has a complexity of 14, somebody's going to use that number to decide where to spend review time. So the rule from day one was: every metric gets checked against an independent tool on real open-source code, and every difference gets explained or fixed. Not "close enough". Explained.
I expected that to be a chore. I figured I'd spend a day documenting the other tools' quirks.
That's not what happened.
What lizard found
First comparison: cyclomatic complexity against lizard on Catch2. 1,381 functions matched, and lizard found 315 that Tezcatl didn't.
The easy story was "header-only templates nothing includes". It was true for 121 of them. The other 194 were in files Tezcatl had parsed. Every member function of every class template was missing.
The cause was one flag. Catch2 builds as C++14 from an MSVC compilation database, and in clang-cl mode before C++20, template bodies aren't parsed until something instantiates them. That's how MSVC does it. Tezcatl's own test fixtures were C++20, where that behaviour is off, so no test of mine could ever have seen it. One flag brought 180 functions back.
Then the reverse: 108 functions only Tezcatl found. 63 of them were = default. My test for "is this a real function" was "does it have a body", and clang writes a body for a defaulted function once it's used. The fixture had one of those too. It passed, because nothing used it.
Then two functions where lizard said 2 and I said 1. A genuine && inside a template that clang couldn't resolve yet, because the library overloads operator&&. It counts now.
After all that, the two tools agree on 99.0% of the functions both find, and every one of the remaining differences has a written reason. In half of them lizard is wrong. In the other half it's a convention I picked on purpose.
Three bugs, all mine, all invisible to a test suite that was green the whole time.
The tests that couldn't fail
That's the uncomfortable part. The suite wasn't just missing cases. Some of the tests couldn't fail at all.
So every rule got a sabotage: break the code on purpose and make sure a test goes red. 13 of the first 14 did. The fourteenth was a percentile calculation, tested on sample sizes where rounding up and rounding to nearest give the same answer. Wrong code passed. Six values tells them apart.
Another one stayed green because of a semicolon. CTest treats a semicolon in a pass pattern as a list separator and passes the test if any piece matches. One of those pieces was a fragment that matched almost any output.
Tezcatl's test suite now covers 95.1% of its own production lines, and there are 142 recorded deliberate faults, each caught by a test. That second number is the one I trust.
The numbers, as they stand
| Check | Result |
|---|---|
| Earthworm (open-source seismic processing), 929 C and C++ files | 0 parse errors, full report in 3.5 s on 16 threads |
| Complexity vs lizard on Earthworm, outside bundled SQLite | 99.0% identical (7,885 of 7,962 functions) |
| Coverage vs gcovr on libmseed, built with gcc 14 | lines, branches and functions identical in all 42 files |
| Coverage vs gcovr on Tezcatl's own code | identical in all 85 files |
The remaining lizard differences on Earthworm trace to code that compiles away: disabled assertions and test macros. Tezcatl measures what the build actually compiles. That's a feature and a limit, and the docs say so.
And one I had to take back
While re-measuring for a capabilities write-up, I rebuilt in Release and Earthworm finished in 3.5 seconds. The number I'd been quoting everywhere was slower.
Every timing in the README, the docs and the release notes had come from a Debug build. The comparisons were fair, Debug against Debug, so the speed-ups held. But the seconds described a build nobody ships, and nothing said which build it was.
A number without its conditions isn't evidence. I've written that down for other people's code. Took me three days to catch it in mine.
How it was built
I built Tezcatl with Claude Code as my pair. That isn't a footnote; it's part of why the checking mattered. An AI pair writes plausible code quickly, and plausible is exactly what a comparison against lizard and a pile of deliberate faults are there to test. Every bug in this post got past both of us and a green test suite. None of them got past a second tool and a deliberate fault.
What it doesn't do
It doesn't check MISRA or CERT rules. It doesn't run your tests; it imports the coverage gcov, llvm-cov or lcov already produced. It isn't certified or audited, and it's brand new: first release was 2026-09-26. If you need any of those, it's not the tool yet.
If you've got a C or C++ codebase and you want a baseline before someone starts changing it, it's free. Tell me where it's wrong.




