How much information do LLMs really memorize? Now we know, thanks to Meta, Google, Nvidia and Cornell
venturebeat.comPublished: 6/5/2025
Summary
While large language models like ChatGPT are trained on vast datasets, their knowledge is a blend of memorization and generalization, relying on patterns within the data rather than rote memory. A recent study reveals that these models can store approximately 3.6 bits per parameter, achieved through random bitstrings to test memorization without external data structures. As model sizes grow, this capacity expands significantly, influencing legal considerations for AI safety and ethical use while opening new avenues for evaluating transparency in AI development.