Five Upcoming Academic Books on Data Work, Big Tech, and Misinformation
Our most hotly anticipated books for fall!
Posted on in Blog Posts
Posted on June 11, 2025 in Blog Posts
Authors:
Rachel Hendrick
Thank you again to everyone who attended the May 22nd webinar, “Equipping Librarians with AI: New Tools to Enhance Workflows.” If you missed the live event, you can always catch up with the recording. Or check out the list of resources Gary and I put together. (Click on the little arrows next to the text to open each section.)
Gary and I love chatting about generative AI tools, and we were sorry we didn’t get to all the questions. To remedy that, I went through the chat transcript to pull out a few more.
One of the best uses for generative AI is as a research assistant. It’s a great way to explore a new topic more deeply and make connections that traditional search engines simply cannot make.
A very basic example of this is the Google AI Overview. (That said, this feature famously advised users to glue cheese to their pizza so it wouldn’t fall off.)
I love to use Perplexity.ai to start researching a new topic. I use the academic dataset and the multi-agent deep seek options. I still read the footnotes to build a historiography, but Perplexity helps me get into the literature more quickly.
But AI for discovery is only as good as the dataset. This is the sweet spot for library vendors. Web of Science, Scopus, and other large scholarly databases have benefited from injecting a little AI into the discovery layer. These AI layers, while expensive, are more complex than their free counterparts. They assist researchers in rewriting prompts, analyzing the suitability of a resource, and identifying gaps in the literature.
AI research assistants are still relatively new, and AI-assisted discovery layers from the big library vendors can be very expensive. But I think this is one of the most useful applications of generative AI in our profession and a great way to introduce students to the benefits of the technology.
🔥 Stay up-to-date with LibTech Insights by signing up for our free newsletter. Just one weekly email with our new blog posts, top tech news stories, and other bonus content. Check out some posts from our archive:
Copyright and AI is a rat’s nest, especially for those of us who are not copyright lawyers. In general, I would advocate for extreme caution in uploading articles and book chapters to a generative AI tool that scans these documents to provide a synopsis or allows a user to chat with the document. However, these tools can really speed up the creation of a literature review. Taking a minute to get to know the nuts and bolts of your document reader can save you a lot of potential heartache.
The first stop, of course, is the tool privacy statement. NotebookLM, for example, says they do not use your personal information nor your document for training. I’m pretty skeptical when it comes to these privacy notices, so I would err on the side of caution and not upload anything that wasn’t freely available on the web (or that I didn’t own the copyright to). Privacy statements can change without warning. So, make this your first stop before you upload a document.
But what about open-access (OA) material? CC-BY licenses are extraordinarily permissive but were developed before generative AI. Again, I’m not a copyright lawyer, but I would proceed with caution. Creative Commons provides some direction on this issue. Short answer? It’s complicated.
Semantic Scholar’s Semantic Reader presents an elegant solution to this question. Ai2 partners with publishers to provide OA content, thereby mitigating any copyright infringement risk. Semantic Reader is enabled only for Semantic Scholar content. It doesn’t allow users to upload their own documents. Subscription databases from reputable library vendors also mitigate any risk of copyright infringement. These aggregators also have licensing contracts with publishers and should assume the risk, should there be any copyright infringement.
Finally, I think it’s important to bring up the concept of context window (aka context length) here. The context window is the amount of text that the AI uses to produce results. (I hesitate to use the word “read” here.) For example, Perplexity has a context window of 4,000 tokens (one token is around four characters) for each query. You can increase that context window substantially if you upgrade to Perplexity Pro. This means that even if you upload a document to an AI-assisted document reader, the tool may not be ingesting the entire document.
Bottom line: if you have any concern about copyright, don’t upload that document.
AI can seem like magic when you first encounter it. How can it know so much? How is it so smart?? Recently, I’ve been playing around with Ithy.ai, a multi-agent AI. Ithy isn’t some creaky, old ChatGPT. Its authoritative answers to my prompts look like well-researched, highly organized reports. They include a fun title, subheadings, graphs, tables, and other visuals.
But after a few searches, it’s easy to see that generative AI is a one-trick pony. Ithy is good at gathering and arranging information, but it’s not creating knowledge. My middle-school-aged son is doing more in-depth analysis of sources than Ithy ever could.
I think Tyler Austin Harper, in his recent Atlantic article, “What Happens When People Don’t Understand How AI works,” puts it best: “LLMs are impressive probability gadgets that have been fed nearly the entire internet, and produce writing not by thinking but by making statistically informed guesses about which lexical item is likely to follow another.”
When I sit down to write a blog post, I’m synthesizing the information I gathered on a topic. I’m making choices on what to include based on my thesis. I think deeply about how I can connect with my audience. What common cultural elements do we share? How can I exploit those commonalities to get my point across? I have never once thought about the probability of what word comes next.
I’m far too practical to be a philosopher, and I dare not enter the debate about what is human intelligence. But here’s my question: Why do we need to recreate human intelligence in a machine? Isn’t it enough that we have electrified rocks and tricked them into processing immense amounts of data for us? Maybe AI is a new kind of intelligence? And maybe that’s okay.
Special collections are my happy place, and there is so much good work happening at the crossroads of metadata creation and AI. For example, representatives from Northwestern University Libraries were at CNI this past spring and gave an update on their project.
LTI contributor Crissandra George, Digital Collections Manager Librarian at Case Western Reserve University, recently wrote about her experience creating a digital exhibition of the Cleveland Bystander. She and Case Western computer science undergraduate Tharun Viswanathan set up a chatbot to create a more interactive experience. The Bystander Chatbot allows users to dive deep into the page of the Cleveland Bystander. Users could synthesize across the entire collection in ways a keyword search simply cannot.
There are lots of exciting ways to enhance special collections with generative AI, but the usual pitfalls of the technology still apply. Is AI introducing bias into how users experience the collection? What are the copyright issues? Is it clear to users that AI was used to create metadata or other elements?
With AI technology becoming more accessible, I’m sure we’ll see more and more instances of these kinds of integrations. It’s an exciting time to be working with special collections!
Our most hotly anticipated books for fall!
Posted on in Blog Posts
What if work could feel remoralizing?
Posted on in Blog Posts
Insights and best practices for teaching AI literacy to history students
Posted on in Blog Posts
What our micro-course participants had to say about AI in libraries
Posted on in Blog Posts