Enterprise-grade version of Local-SLM-Data-Cleaner: a small language model (SLM), fine-tuned entirely on synthetic data, that normalizes unclean SAP-style master data to a documented house convention. It is built to run in secure, air-gapped enterprise environments.
Where the original repo is a beginner-friendly demo for a Mac laptop, this version targets production deployments: editable client-specific convention specs, a (manual) review queue with an append-only audit trail, containerized offline serving with vendored model weights, and eval-gated releases.
Every company running SAP knows the problem: the same supplier exists three times with three spellings, countries are written as "Germany", "Deutschland" or "DE", dates and amounts come in German and US formats, and missing values are encoded in five different ways. Cleaning this by hand is slow, as writing rules for every possible mistake is tiresome.
This project takes a different route. Your data standard is written down once, as a readable document. From that standard, the system invents thousands of practice examples (dirty record IN, clean record OUT) and teaches a small AI model the standard. The result is a model that cleans records the way your rules would, and also handles the misspellings and odd formats no rule anticipated.
The decisive point: your data never leaves your machine. There is no cloud service, no API subscription, no data transfer to anyone. The model is a file on your own hardware, small enough to run on an ordinary office machine, and it works without Wi-Fi, or the network cable.
Sensible Stammdaten an eine ausländische Cloud-KI zu senden, ist unter der DSGVO oft keine Option, und für viele Häuser auch keine Frage der Rechtslage, sondern der Grundhaltung. Dieses Projekt zeigt den anderen Weg:
- Alles läuft lokal. Das Modell ist eine Datei auf Ihrer eigenen Hardware. Kein Cloud-Dienst, kein Abo, keine Token-Kosten, keine Datenweitergabe. Auf Wunsch komplett vom Netz getrennt (Air-Gap), nachweisbar durch die Container-Konfiguration, nicht nur durch Versprechen.
- Trainiert ausschließlich auf erfundenen Daten. Kein einziger echter Datensatz wird für das Training verwendet. Ihre Daten kommen erst zur Laufzeit ins Spiel und bleiben im Haus.
- Jede Änderung ist nachvollziehbar. Ein unveränderliches Protokoll (Audit-Log) hält für jeden Datensatz fest, was geändert wurde, mit welcher Konventionsversion und mit welchem Modellstand. Unsichere Fälle gehen in eine manuelle Prüfschlange, nie stillschweigend durch.
- Das Basismodell ist austauschbar. Wer kein chinesisches Modell einsetzen möchte, wählt ein europäisches (z. B. Ministral-8B oder Mistral Nemo von Mistral AI in Frankreich, Teuken-7B von Fraunhofer, gefördert vom BMWK, oder EuroLLM aus einem EU-Projekt) oder ein US-Modell mit MIT-Lizenz. Details in der Tabelle unten.
Beratung und Umsetzung: mbitai.com.
-
The rules. Your data standard (allowed country codes, legal forms, date formats, how missing values are written) lives in one readable file client-specific: conventions/default.yaml. A data steward can open and edit it; no programming involved. Switching clients means switching files, not rewriting software.
-
Practice data, invented from the rules. The system generates thousands of fake records (invented company names, invented IBANs), makes them dirty in realistic ways, and computes the correct cleaned version from the rulebook. No real data is used at any point, which is why the training itself is GDPR-uncritical.
-
Training, on ordinary hardware. A small open-source model (about 1 GB; see the model table below) practices on those examples until it reliably produces clean records. This runs on a normal laptop in under an hour; no GPU cluster, no cloud.
-
Operation with a double check. In production, every record the model cleans is immediately re-checked against the rulebook. If the answer violates any rule, the deterministic rule engine takes over as a safety net. The model adds value on the messy edge cases; the rules guarantee a floor it can never fall below.
-
Nothing uncertain slips through. Records the system is not sure about (rule violations, low model confidence) go to a manual review queue. A person decides; the decision is recorded.
make reviewlists what is waiting. -
Everything is on the record. Every cleaning decision is appended to an audit log: input, output, every single change, confidence, the exact version (hash) of the model weights and of the convention file. The log is append-only: entries are never edited or deleted, resolutions are added as new entries. This is what your auditor and your Datenschutzbeauftragter will ask for.
-
Sealed delivery. For production, everything ships as one container that is run with its network stack removed (
--network none). The container refuses to start if the model weights do not match the fingerprint pinned in version control, so nobody can swap the model unnoticed. Records go in and out through a mounted folder only. See deploy/README.md. -
A quality gate on every change. The repository carries a suite of pinned test cases, including adversarial ones (is "Bavaria" wrongly "corrected" to a country? is "mbH" recognized as GmbH? is a US date misread as a German one?). Continuous integration blocks any change that alters this documented behavior. Releases are gated by measured accuracy, not by hope. Pinned gold fixtures (
fixtures/gold.jsonl) are separate from generateddata/test.jsonland are never overwritten bymake data. -
Semantic alias resolution (optional). An embedding soft-normalizer can sit in the cleaning pipeline to catch near-misses the alias map never listed. Off by default; see Semantic embedding layer below.
-
Eval hardening (oracle peer, balance, privacy). Every serious report compares oracle (
normalize_record) · base SLM · fine-tuned SLM on the same fixtures. Training refuses to start if the synthetic train set fails the coverage balance gate. Real extracts stay underfixtures/real/local/(gitignored); see docs/PRIVACY.md. Stewards edit conventions/ only; no code changes.
Every controlled-vocabulary field (country, legal form, status, currency, unit) is normalized in a fixed order. Later stages only run when earlier ones miss. This way, documented aliases stay authoritative and the optional embedding layer never overrides a known mapping:
raw value
→ blank / sentinel? → null
→ exact controlled set? → keep
→ deterministic alias? → canonical (conventions/*.yaml)
→ embedding near-miss? → canonical (optional, USE_EMBEDDINGS=1)
→ otherwise → as-is (grounding: flagged, not invented)
Why this order. Putting embeddings after the alias map preserves the
house standard: "germany" → DE always comes from the YAML, never from a
similarity guess. Embeddings only help with novel paraphrases
("Nederlands", "Federal Republic of Germany") that nobody thought to list.
Grounding guarantee. Values below the per-field similarity threshold
pass through unchanged. Regions and true unknowns must not be "helpfully"
corrected: Bavaria stays Bavaria, Atlantis stays Atlantis, the typo
GmbbH stays GmbbH. The adversarial suite pins these cases.
Labels stay honest. Synthetic near-miss corruption runs only when
USE_EMBEDDINGS=1, so normalize_record() can reverse it. Without the flag,
training data never teaches the model to pass through a value that production
would clean.
Activate with USE_EMBEDDINGS=1 (off by default; no overhead when unset).
pip install sentence-transformers>=3.0
make embed # download + cache the embedding model
USE_EMBEDDINGS=1 make data # near-miss examples get correct labels
USE_EMBEDDINGS=1 make eval-gate # include semantic_alias adversarial cases
USE_EMBEDDINGS=1 make demo # use embeddings at inferenceThresholds are per field under embedding_thresholds: in
conventions/default.yaml. The resolver compares
the unknown string against canonical codes and every alias key (readable
anchors). Matching bare ISO codes like DE alone is not enough: short codes
carry almost no semantics for an embedding model.
| Model | Size | Role | Notes |
|---|---|---|---|
| BAAI/bge-m3 | ~570M | Default in core/embedding_lookup.py |
Multilingual (100+ languages), strong on short phrases, loads via sentence-transformers. ~2.2 GB download, ~1.2 GB RAM at inference. |
bge-m3 is the right default for this stack: it runs on CPU or Apple Silicon
MPS, needs no NVIDIA GPU, and handles the German/English/French mix typical
of European master data. To swap the encoder, change the model id in
core/embedding_lookup.py, retune embedding_thresholds:, and re-run
make embed plus the adversarial suite before shipping.
The demo default is Qwen3-0.6B, an Apache-2.0 open-source model originally published by Alibaba. Two frank statements about that:
Objectively, the origin of the weights is not a data-protection issue in this architecture. The weights are a static file of numbers, fine-tuned by us on synthetic data, running in a container that has no network. There is no channel for data to flow anywhere, regardless of who trained the base model.
Commercially and politically, the choice still matters. Procurement, works councils and security officers have preferences, and "the AI is Chinese" can end a conversation in a German enterprise regardless of the technical facts. This stack is therefore deliberately model-agnostic: the convention specs, synthetic data, eval suite, audit trail and container do not care which base model sits inside! Swapping means retraining (cheap, as the training data is already generated) and re-running the same eval gate.
| Base model | Size | Origin | License | Notes |
|---|---|---|---|---|
| Qwen3-0.6B | 0.6B | Alibaba (CN) | Apache-2.0 | Demo default: smallest, fastest, runs on 8 GB laptops |
| Teuken-7B-instruct | 7B | Fraunhofer / OpenGPT-X (DE) | Apache-2.0 | German-made, BMWK-funded, all 24 EU languages. Nice sovereignty pick, but needs a bigger training machine (≥32 GB or a GPU) |
| EuroLLM-1.7B | 1.7B | EU project | Apache-2.0 | European, all official EU languages, still small footprint |
| Ministral-8B-Instruct-2410 | 8B | Mistral AI (FR) | Apache-2.0 | French-made edge model; LoRA-trains on a 16 GB Mac (~4.5 GB at 4-bit). Multilingual incl. DE. Strong EU pick when 0.6B is too small but Teuken is too heavy |
| Mistral Nemo Instruct | 12B | Mistral AI (FR) | Apache-2.0 | Mistral+NVIDIA collab; better instruct quality than the 7–8B class. Plan for ≥16 GB RAM at 4-bit (~7 GB weights) |
| Phi-4-mini | 3.8B | Microsoft (US) | MIT | Most permissive license of the field; strong quality, mid-sized |
| Gemma 3 1B | 1B | Google (US) | Gemma Terms | Capable, but not pure open source: Google's terms attach a use policy |
| SmolLM3-3B | 3B | Hugging Face (US/FR) | Apache-2.0 | Fully documented training recipe, strong transparency story |
Practical recommendation for EU customers:
- EuroLLM-1.7B: smallest European footprint; still fits the 8 GB demo laptop.
- Ministral-8B-Instruct-2410: the sweet spot for made in the EU without a server room: French company, Apache-2.0, trains on an ordinary 16 GB Mac, and enough capacity for messy German master-data edge cases that a 0.6B model struggles with.
- Mistral Nemo 12B: same sovereignty story with more headroom; pick it when you have ≥16 GB RAM and want better zero-shot quality before fine-tuning.
- Teuken-7B: where made in Germany carries the most weight and a bigger training machine (≥32 GB or a GPU) is available.
Phi-4-mini remains the US quality/license fallback if procurement rules out European weights entirely.
To swap, rerun the pipeline with a different MODEL_PRESET. Run make clean
first so adapters and GGUF files from the previous model are not reused:
make list-models # show available model presets
make MODEL_PRESET=minicpm5-1b clean model data train fuse gguf serve eval
# All pipeline commands after switch respect the preset:
make MODEL_PRESET=minicpm5-1b train # re-run training only
make MODEL_PRESET=minicpm5-1b eval # re-run evaluationOr set it once for the session:
export MODEL_PRESET=minicpm5-1b
make clean model data train fuse gguf serve evalAvailable presets:
| Preset | Base model | Size | GGUF repo (baseline) | Quant | ALIAS |
|---|---|---|---|---|---|
qwen3-0.6b |
Qwen3-0.6B | 0.6B | Qwen/Qwen3-0.6B-GGUF |
Q8_0 |
qwen3-0.6b-cleaner |
qwen3.5-0.8b |
Qwen3.5-0.8B | 0.8B | unsloth/Qwen3.5-0.8B-GGUF (community; no official Qwen GGUF yet) |
Q8_0 |
qwen3.5-0.8b-cleaner |
minicpm5-1b |
MiniCPM5-1B | 1B | openbmb/MiniCPM5-1B-GGUF |
Q4_K_M |
minicpm5-1b-cleaner |
Each preset sets four variables together. You can still override any individually:
make model train fuse gguf MODEL_PRESET=minicpm5-1b GGUF_QUANT=Q8_0(One caveat: each architecture must be supported by the MLX trainer and llama.cpp! Qwen, Gemma, Phi, Llama-family, SmolLM, MiniCPM, and Mistral (including Ministral and Nemo) are all supported today.)
core/ the convention engine (loads the YAML spec, single source of truth)
core/embedding_lookup.py (optional) post-alias soft-normalizer
core/llama_client.py (shared live-model HTTP client)
(default encoder: BAAI/bge-m3; swap-friendly)
conventions/ editable client-specific convention specs (YAML), product surface
see conventions/README.md
fixtures/ pinned gold + unseen-noise holdout (never written by synth/)
fixtures/real/local/ gitignored real extracts
synth/ synthetic messy->clean data generator → data/ only
eval/ eval harness + pinned adversarial suite
(legal forms, formats, grounding, semantic_alias)
runtime/ clean service: model -> validate -> algorithm safety net + audit + review
--min-confidence (default 0.9) → audit/review-queue.jsonl
train/ check_balance.py (train gate); use `make train/fuse/gguf/serve`
reports/ markdown eval reports (oracle · base · FT)
docs/ PRIVACY.md and related
deploy/ air-gapped container: Containerfile, entrypoint, security notes
.github/ CI eval gate: regressions cannot merge
Makefile every pipeline step as `make <command>` (see `make help`)| Piece | Command / path | Role |
|---|---|---|
| Steward YAML | conventions/ + README |
Edit rules without code |
| Sacred gold | fixtures/gold.jsonl |
Hand-pinned; make oracle |
| Unseen holdout | fixtures/holdout_unseen_noise.jsonl |
OCR-ish family excluded from train |
| Oracle peer | make oracle / make report-oracle |
Permanent rules row in reports |
| Balance gate | make check-balance (also prerequisite of make train) |
Fail closed on thin coverage |
| Privacy | make privacy-check / docs/PRIVACY.md |
No real extracts in git |
| Review queue | runtime/clean.py --min-confidence 0.9 |
Low confidence never silent-accept |
| Real smoke | fixtures/real/README.md | 20 local lines before production FT claims |
Required report shape: reports/template-eval.md.
make fixtures # rebuild gold + unseen (commit only when intentional)
make report-oracle # oracle on gold → reports/oracle-gold.md
make data check-balance
make eval-gate # privacy + gold + synth test + adversarial + unseen
# optional: make oracle ORACLE_DATA=fixtures/holdout_unseen_noise.jsonlmake setup # install Python deps + mlx-lm
make fixtures # ensure pinned gold + unseen holdout exist
make data # generate synthetic train/valid/test under data/
make check-balance # coverage gate (also runs before make train)
make sanity # verify generated test vs oracle (~100%)
make eval-gate # full regression gate
make train # LoRA fine-tune (blocked if balance fails)
make fuse gguf # package the trained model for serving
make pin-model # vendor the weights + pin their hash for the container
make serve # serve it locally...
make eval # score fine-tuned model on fixtures/gold.jsonl
make review # list records waiting for manual review
# optional: embedding soft-normalizer (raises score on near-miss aliases)
pip install sentence-transformers>=3.0
make embed
USE_EMBEDDINGS=1 make data eval-gateClient-specific convention: make data CONVENTION=conventions/<client>.yaml
For the full step-by-step walkthrough and the concepts behind the approach (knowledge distillation, LoRA, quantization, grammar-constrained decoding), see the origin repo's README.
Built on Local-SLM-Data-Cleaner by mbitai, which remains the public demo and tutorial for this approach. All sample data in both repos is synthetic and invented; no client data is involved at any point.
AGPL-3.0 (see LICENSE).
For commercial licensing without AGPL obligations, or help applying this to your own master data migration or data-quality work, contact mbitai.com.