Model Genome: Fingerprinting Whether an LLM Was Trained From Scratch or Derived
Outsiders can assess whether a foundation model was truly built from scratch by analyzing architecture configurations, tokenizer overlap, and weight embeddings using a reproducible fingerprinting pipeline. While architecture and tokenizer artifacts provide the strongest evidence, weight analysis has limitations and cannot cleanly distinguish continued pretraining from training from scratch.
https://huggingface.co/blog/mayafree/model-dna
huggingface.co
Model Genome: Fingerprinting Whether an LLM Was Trained From Scratch or Derived
A Blog post by Proto_AGI on Hugging Face
August 18, 2026 135 1