“Vintage” large language models (LLMs) are LLMs trained on historical material, which, absent of context, replicate past prejudices and biases without critique for modern audiences. This replication could potentially reform past prejudices in untrained audiences. I had the opportunity, with my colleagues Jessica Jack and Jacob Polay of the University of Saskatchewan’s History department, to publish a piece in Active History. The piece is a warning on the negative impacts that vintage language models like talkie can have when presented without context. Beyond the scope of the article we proposed solutions and improvements to these vintage models for future work. These include: making the historical training data public and developing retrieval augmented generation (RAG) models rather than chatbots. RAGs can provide references to specific records, as opposed to chatbots that synthesize a of a morass of training data and present that data as a decontextualized whole.
The full post can be found here.

