OLMoTrace: An AI That’s Transparent about Its Training Data

A LLM without a black box

A librarian using OLMoTrace for transparency about its training data

It’s no secret that AI companies require vast datasets to train their generative AI models. What is almost always a secret is what these datasets are. (Widespread copyright infringement might be a contributing factor to this secrecy.) Some AI companies, like Perplexity, have made it their selling point to offer users transparency into their models’ process for generating responses.

In this installment of ✨LibTech Tools✨, Choice editor and publisher Rachel Hendrick and information services industry expert Gary Price cover a new AI model that is making great advances in openness of both its model and its training data, Ai2’s OLMoTrace. (View Ai2’s blog for additional background.) Gary and Rachel give a quick tutorial of the model and highlight its pedagogical value.

🚀 PS: Hit that subscribe button on YouTube!