Mojibake: Zero-Dependency C Library for Unicode 17
17 Jul 2026
A new open-source library called Mojibake surfaced this week via a Show HN post, pitching itself as a zero-dependency, self-contained solution for Unicode 17 text processing. Written in C11 and compatible with C++17, it targets developers who need normalization, bidirectional text, collation, and other Unicode-standard algorithms without pulling in heavyweight dependencies like ICU.
What Mojibake offers
According to the project's announcement, Mojibake ships as a downloadable amalgamation package — just two files, mojibake.c and mojibake.h — making it straightforward to drop into an existing codebase. It's released under the MIT License, which typically allows flexible commercial use without licensing overhead.
On the feature side, the library covers a broad swath of Unicode functionality:
- Normalization (NFC/NFD/NFKC/NFKD)
- Case conversion
- Grapheme cluster segmentation
- Bidirectional text handling
- Collation
- Confusable detection
It accepts and outputs UTF-8, UTF-16LE, UTF-16BE, UTF-32LE, and UTF-32BE encodings, and the maintainers say it runs across Linux, macOS, FreeBSD, OpenBSD, NetBSD, and Windows 10/11.
The numbers behind the claims
The project reports over 1.5 million assertions run against it, including official Unicode conformance suites. It's also been fuzzed with libFuzzer and is described as AddressSanitizer and UBSan clean — engineering signals often cited to demonstrate baseline robustness in low-level C code.
One practical detail for size-conscious builds: disabling the MJB_FEATURE_CHARACTER_NAMES feature reportedly reduces binary size by approximately 30%, a potentially useful lever for embedded or resource-constrained deployments.
What's missing from the picture
The report accompanying this launch flags several gaps. There's no stated version number, release date, or information on project maturity. Nothing is disclosed about team size or maintainer background, and there are no performance benchmarks comparing Mojibake to established libraries like ICU. Adoption data — downloads, production usage, or case studies — is also absent, as is any indication of how the 1.5M+ assertions map to actual Unicode conformance test coverage percentages. No roadmap for future features has been shared either.
Risks to weigh
As a new and unproven library, Mojibake may lack the battle-testing and edge-case handling that comes from years of real-world use in established Unicode libraries. Zero-dependency, low-level C libraries can also carry higher integration and memory-safety risk if they haven't been widely audited — fuzzing claims aside. And with no stated team size, there's an implicit bus-factor risk: reliance on a single maintainer or small team could threaten long-term support.
Why founders should care
For early-stage teams building text-processing features — search, internationalization, form validation, or messaging products — Mojibake's zero-dependency, MIT-licensed design may lower integration friction compared to bulkier alternatives. Cross-platform support could plausibly suit startups deploying across varied OS environments without extra porting work, and the option to disable features like character names for a ~30% binary size reduction could be attractive for embedded or constrained platforms.
That said, the lack of adoption data, performance benchmarks, and maintainer transparency means founders should likely treat this as an early-stage bet rather than a drop-in replacement for proven infrastructure. Teams considering it for production would be well advised to run their own benchmarking and stress-testing before committing, particularly for mission-critical text processing where edge-case failures could be costly.