Five Upcoming Academic Books on Data Work, Big Tech, and Misinformation
Our most hotly anticipated books for fall!
Posted on in Blog Posts
Posted on April 16, 2025 in Blog Posts
Authors:
Rachel Hendrick
This year’s Coalition for Networked Information (CNI) marked a meaningful conclusion to Cliff Lynch’s tenure as the Executive Director of CNI. Cliff’s legacy lies in creating excitement in the profession about the future of technology in libraries. For those who will miss Cliff’s yearly keynote address, take heart that a Festschrift is in the works, available in the July issue of Libraries and the Academy and published in open access via Project Muse.
As usual, the CNI team put together a program that carefully balanced the current state of networked information with just enough futurology to keep us all thinking. If you couldn’t make it to Milwaukee last week and you’ve already read the Meeting Roadmap, this post will offer a few takeaways and some reading to bring you up to speed.
Training generative AI models has long been an opaque process and the subject of a few copyright disputes. Academic libraries increasingly face attacks from bots harvesting delicious metadata to train large language models (LLM). And yet generative AI models still display implicit bias.
The Authors Alliance gave an update on their initiative to create better training documents for LLM. The idea is that by curating the training documents, we can control for implicit bias in generative AI. This Public Interest Corpus will pull from libraries, archives, and other “memory organizations” to create a multimedia database of open-access items that can create better LLMs.
To combat web crawling, the Internet Engineering Task Force (IETF) AI Preferences Working Group (AIPREF) has proposed new guidelines that will allow website publishers to mark their sites as off-limits to data harvesting bots. These improved guidelines for robots.txt files may help libraries manage the web crawler traffic that has, on occasion, shut down OPACs.
Further reading:
🔥 Stay up-to-date with LibTech Insights by signing up for our free newsletter. Just one weekly email with our new blog posts, top tech news stories, and other bonus content. Check out some posts from our archive:
Another hot topic was metadata and how we can build infrastructure for better collaboration and data security. Representatives from the Bue Core project, which aims to create a community-owned linked data environment, gave an update on their project. In 2025, this project moved into prototype, and the team hopes to have a viable product by 2026.
A representative from the Collaborate Metadata Enrichment Taskforce (COMET) gave an update on their work to create persistent identifiers (PID). The hope is that this project will create a scalable approach to better metadata collaboration via community-led governance standards.
Further reading:
In other AI news, two libraries demonstrated how they use LLMs to enhance the discovery of their collections. The technology team at the Baker Library at Harvard Business School gave an update on their knowledge platform. They outlined their dynamic knowledge graph and spoke about the importance of these data models in rethinking AI-enhanced discovery layers.
Northwestern University also gave an update on their Institute of Museum and Library Studies (IMLS) grant to create an AI discovery layer in their special collections. The presentation included a tour of the chat-based search option.
Further reading:
Our most hotly anticipated books for fall!
Posted on in Blog Posts
What if work could feel remoralizing?
Posted on in Blog Posts
Insights and best practices for teaching AI literacy to history students
Posted on in Blog Posts
What our micro-course participants had to say about AI in libraries
Posted on in Blog Posts