5 Questions to Ask Your Vendors about AI

Cut through the hype and get real answers about AI products

A librarian asking questions to a vendor about AI tools

Generative AI is quickly becoming ubiquitous. Its integration is not just limited to our personal technology, but now also includes library resources, with vendors constantly introducing new tools or integrations into existing platforms. Meanwhile, there are dozens of emerging AI platforms that we must measure against our library’s tried-and-true resources to gauge which tools truly bring a new perspective or added value to our collections. 

It’s an overwhelming time to vet and license library resources. To cut through the noise, I’ve narrowed down the many things I want to ask vendors to five questions to better understand how a new AI tool might bring new depth or insight to my library’s user experience.

1. What content is your tool trained on?

The future of generative AI models as we know them will heavily depend on whether the use of copyrighted content to train models violates copyright or falls under Fair Use. It will take years for the copyright-and-AI question to be decided, but if the courts ultimately rule that all content used to train models must be licensed, many existing models could become unusable. 

While this unfolds in court, I am interested in ethical data acquisition by my vendors. AI tools typically follow one of three training approaches: training a full model, supplying a full model with a specific knowledge base, or training a small language model. Each approach to an AI model has its pros and cons. Large language models respond more predictably but aren’t trained as well for specific purposes, whereas small language models can focus on specific purposes but potentially sacrifice accuracy.

The vendor’s response to this question also provides insight into the ethics of their model. Ideally, all content should be properly licensed and clearly attributed. If user inputs or uploads are used for training, vendors should clearly notify users and warn them about possible copyright infringement. A vendor who is unwilling or unable to address how their model is trained and how they source model content is cause for concern.


🔥 Stay up-to-date with LibTech Insights by signing up for our free newsletter. Just one weekly email with our new blog posts, top tech news stories, and other bonus content. Check out some posts from our archive:

🌟 Register now for our free email micro-course on AI literacy!


2. What sets your tool apart from other vendors or from freely available models like ChatGPT?

While there are many capable generative AI tools on the market, there is a limit to the uniqueness and value that they can offer library users. This is especially true in a world of constantly shrinking budgets where content redundancy no longer makes sense for many libraries. When evaluating newly available AI resources, libraries must carefully consider the incremental benefits of a licensed tool and whether they outweigh the free functionality users can access elsewhere. Whether you love or hate AI, there is an undeniable appeal in a general-purpose tool that can help with a range of everyday tasks. Well-trained models with sufficiently sized corpuses can already help users navigate research tasks, such as literature reviews. These tools may not have direct access to our licensed content, but can help construct meaningful keyword searches, identify relevant databases, and summarize seminal research without specific content.

If the library licenses and funds an AI resource or tool, the tool needs to provide substantial value over the free version of general-purpose models. Kickstarting your research in ChatGPT is a questionable practice from a research and copyright lens, but I guarantee that library users are already engaging in these behaviors. I may disagree with the practice, but I acknowledge that an AI tool that is only usable in a limited capacity is less desirable than a general-purpose model that can be tapped for a variety of uses.

Fundamentally, a library-licensed model needs to bring true novelty to justify adoption. Resources that operate in limited environments, such as a model that can only return information about a single database, run the risk of frustrating users when it is unable to answer their questions completely. Moreover, limited tools may devalue the library’s resources by conveying that library resources are less useful than a free tool.

3. How does your AI tool or platform protect the privacy of users and the intellectual property of other vendors?

Users are already overly comfortable sharing their personal information with AI models. While I won’t say that free models excel at safeguarding user data, users are empowered to select models of their own choosing. Data privacy is rapidly evolving in a post-AI world, but users assume that most of their conversations will never see the light of day and will stay locked behind logins that they set themselves.

Patron privacy is a core value of librarianship. Users trust that resources provisioned by the library protect their best interests and that their data is safeguarded. Any adopted AI models or tools must reflect this trust. Our vendors must provide transparent guidelines on how user data is captured and used, especially if it is being used to retrain their tool. In a perfect world, user data is not captured by a model, remains anonymous, and cannot be used for training (i.e., a closed model). Otherwise, resources should clearly inform users of any saved data and intended future use. 

4. What do librarians and users need to know before engaging with your AI tool or platform?

There is always some level of onboarding for all new library resources. Some teams will have the bandwidth to deep dive into each new resource, while some libraries may only have a few minutes to review a new resource before they begin working with users on that platform. If you cannot arrange a trial for an AI platform, your conversation with the vendor is the best measurement of how painless platform adoption will be for your library. 

AI should streamline or automate something, so platforms that are difficult to navigate, require unintuitive engagement, use unusual jargon, or return unreliable results may be deal-breakers. I am wary of new resources that purport to return excellent results but are difficult to use, because they not only require a time investment from library staff to learn but also result in frustrated users. Platforms don’t have to be perfect out of the box, but I am cautious of a tool that may add minimal value but drastically increase the stress and workload on library teams.

5. What is the vendor’s long-term AI strategy? 

In my experience, few vendors have an iterated multiyear plan for releasing, updating, and expanding their AI tool. The proliferation of AI is not sustainable in more ways than one, with continuous model growth and year-over-year costs being of specific concern to libraries. For example, it is estimated that AI will run out of training data sooner rather than later, which will hinder continuous model improvement. Depending on how an AI tool is constructed and AI copyright law, models may not drastically improve with time or improvement may be slow. If a platform doesn’t meet your expectations now, it may not be able to meet them later, either.

Moreover, if each library vendor platform rolls out an AI tool at an introductory rate or begins to build in AI components that can’t be separated from overall platform costs, these resource costs will quickly outpace library budgeting. Many of my library’s current resources have incorporated AI components at no additional cost. However, these integrations sit inside of the previously existing platform, and few vendors have been able to speak to whether those integrations will continue freely in the coming years. I am also unsure whether vendors will offer a non-AI version of their platforms or if their traditional platforms will continue as we know them at all. It is not a stretch to imagine that some library vendors will pivot to a platform with a non-optional AI component and then charge a premium for their AI-powered resources.

Ultimately, vendors may not have the perfect answer to every question but asking the right questions positions librarians to dig deep and advocate for their communities. AI will continue to evolve at a pace that libraries cannot always match. However, approaching these tools with informed curiosity allows us to separate genuine innovation from fleeting hype and ensure that the resources we license continue to serve users ethically, responsibly, and well.