How much information do LLMs really memorize? Now we know, thanks to Meta, Google, Nvidia and Cornell

venturebeat.comPublished: 6/5/2025

Summary

While large language models like ChatGPT are trained on vast datasets, their knowledge is a blend of memorization and generalization, relying on patterns within the data rather than rote memory. A recent study reveals that these models can store approximately 3.6 bits per parameter, achieved through random bitstrings to test memorization without external data structures. As model sizes grow, this capacity expands significantly, influencing legal considerations for AI safety and ethical use while opening new avenues for evaluating transparency in AI development.