Week 5: Evaluating AI

AI literacy essentials for academic libraries
Choice and Clarivate have teamed up to develop an eight-week newsletter-based course on generative AI literacy for academic library workers. If you are AI curious—and short on time—then this course is for you!
Each week contains three sections:
- Introduction: An overview of a core AI literacy competency
- In Practice: An interview or case study
- Reading Room: Further reading material focused on the topic
- The module closes with a quiz and discussion questions.
Make sure you’re registered to receive all eight weeks directly in your inbox! Know someone who might be interested in this micro-course? Share the registration page with your colleagues.
In week five, we look at how to evaluate AI resources and output.
Table of Contents
Introduction
Evaluating AI Resources for Scholarly and Library Workflows
by Juan Denzer, Engineering and Computer Science Librarian at Syracuse University Library
Librarians may approach the evaluation of generative AI applications the same way they approach the evaluation of any digital tool. User needs, functionality, quality of content, cost, and privacy considerations do not vary. However, generative AI tools present new evaluation challenges. Ethical and legal factors, performance issues (e.g., hallucinations), technology sustainability, and vendor reputation must be carefully considered when evaluating an AI tool for your library.
Before evaluating an AI research tool, it is important to take it on a test drive. Library vendors and publishers are more open to providing institutional trials, while other developers might only offer limited individual trials. It is also important to provide trials and get feedback from both patrons and library professionals from your institution.
User Needs and Functionality
User needs and expectations are evolving almost as quickly as generative AI. Users expect results that are instant and easy to read and understand. Make sure that the tools you evaluate meet those needs. The functionality of these tools should be simple and easy to navigate.
Evaluate if the generative AI tool is actually adding value to your collection. If the AI layer is simply a gimmick or does not add any meaningful functionality, then it might not be a tool worth recommending. The exception is AI research assistants for existing content collections. These tools will have the similarities on the front end in terms of search and discovery, but under the hood, AI enhances almost all aspects of discovery.
Content and Quality
Evaluate both the large language model (LLM) and the retrieval augmented generation (RAG) layers to get a better sense of how the AI tool is producing content. Most LLMs are trained on public information from websites and online forums. This includes open access scholarly resources, but also websites such as Reddit. The real value library database vendors provide is developing RAG layers curated for scholarly workflows. Although the LLM engine may be the same as a free chatbot, an AI tool with a RAG layer designed specifically for research will be much more useful and produce far fewer hallucinations.
Context window or token length is another important metric when evaluating the nuts and bolts of an AI tool. This is the amount of the document the LLM evaluates to generate results. A larger context window or longer token usage takes more processing power, but it also produces better results. For example, GPT-3.5 evaluates up to 4,096 tokens, while GPT-4 can evaluate up to 128,000 tokens. (A token is about 4 characters and the average word length in English is 4.7 characters, per Google’s AI overview.) Understanding not only the information the LLM is evaluating (the RAG layer), but how much the LLM is evaluating (the context window or token length) will empower you to make better decisions about the best AI tool for your collection.
When evaluating a tool, it is best to try as many different queries as possible. Remember to think like a patron who is both experienced and inexperienced when using research tools. The goal is not to try and stump the AI or get a hallucinated result. If you do find the system is generating hallucinations, evaluate and rank them. Keep track of those hallucinations and reach out to the AI provider and ask about the quality of the results.
Sustainability and Reputation
Sustainability, longevity, and maintainability are all key factors in evaluating the total cost of a generative AI tool. Consider whether the tool is open source and supported by a development community, or created and maintained by a private company. Is a free open-source tool worth the effort and do you have the IT resources to support it? Or is a proprietary product maintained by a private publishing company able to provide better value? What about the reputation of the company and the tool? Think about who is developing the software. Is it a company that is known for developing library and academic products? Perhaps they are a startup with little to no academic experience. Look at the company’s creators and development team.
Pricing
As with other collection resources, the pricing structure of an AI tool will vary. For example, some tools offer individual or institutional subscriptions. Others sell credits that may limit heavy usage of AI features.
Consider the size of the LLM being used. Some developers might use smaller LLMs that consume less energy. These savings might be passed along to users and make the cost of the tool cheaper. Make sure to ask if the AI tool uses a large LLM or a smaller LLM. Also, ask if the LLM is an in-house proprietary LLM or a subscription to a larger company such OpenAI’s GPT or Google’s Gemini.
Privacy and Legal Considerations
Privacy should always be considered when evaluating the tool. Make sure to research not only the privacy policy of the AI tool but the privacy policy of the LLM used. Ensure that the content that a user provides is protected. If user data is used for training, make sure users are immediately notified at each session. Make sure data privacy policies cover all your users; some LLMs have separate policies for different regions in the world. Some tools allow users to upload documents, which has the potential to violate copyright laws. Be sure to evaluate the legal liability scenarios that may arise from these kinds of tools.
Performance and Limitations
Accuracy, latency, scalability, ethical, and human evaluation are all ways the performance of a generative AI tool may be evaluated. How accurate are the results of a query? Latency measures how long it takes the tool to respond. Although LMM benchmarks can be found in the literature, it is best to use your own metrics to test latency. Scalability measures how well a tool will handle the increase in users. This is especially important when considering the number of users for your library. Ethical considerations include bias and harmful content. For example, are the results biased compared to traditional research searches? Consider the limitations of the AI tool. Is it capable of handling other forms of media beyond text?
Licensing and Agreement
The L&A (License and Agreement) is what brings all this together. Be sure to consult the notes you created during your evaluation of the tool as you negotiate your agreement. Be sure that your agreement considers user and data privacy, performance and limitations, and legal considerations.
In Practice
How SUNY Empire State University Library Evaluates AI for Collection Development and Curricula
As an online-only institution, digital accessibility and innovation are the cornerstones of SUNY Empire State University. Serving adult and non-traditional learners, SUNY Empire is also an Autism Supportive College, making inclusive and reliable resources all the more essential. The introduction of artificial intelligence (AI) is quickly influencing student and faculty needs. Now, SUNY Empire Library must determine how to best evaluate AI tools for collection development, curricula, and research.
Click here to read more about how SUNY Empire State University library is building their collection of generative AI resources.
Addressing Change Fatigue
“Empire is in a similar place as most other US Higher Education institutions: there are many diffuse AI activities, and different groups are working on AI from different perspectives,” says Shannon Pritting, Director of Library Services, Open and Digital Learning Assets at Empire Library. “With this in mind, the library has realized that we need something simple that can adjust as we develop policies and processes on AI tools in teaching, learning, and research activities.” Flexibility and simplicity are vital due to rapid AI development and what Pritting refers to as “change fatigue” across campus. Pritting underscores, “Keeping things simple has served us well and has allowed us not to contribute to AI change fatigue by not involving too many stakeholders too often while putting in simple safeguards so we don’t implement something that will be disruptive.” In addition to the need for AI guardrails, usability, patron privacy, and mitigating bias are concerns as collection development expands to include AI tool licensing.
Drawing on a Flexible AI Framework
Thus, the need for an evaluative framework to assess generative AI tools. With a core focus on “the impact the AI tool will have on teaching, learning, or research,” Empire Library turned to the EDUCAUSE Framework for AI Literacy. “Although this framework isn’t intended to be used as an evaluative framework, it was useful to determine when we should wait and engage faculty and others in depth before implementing a tool, and when we can enable the AI tool without widespread input,” Pritting explains. Empire Library also joined the Ithaka S+R cohort on defining and implementing AI Literacy. Involvement in the cohort aims to help better understand student and faculty expectations and uncover opportunities for new product offerings.
A Collaborative Effort
“Whenever we’re adopting educational technology, we use a team-based approach to look at how the technology will impact all learners and how it changes student research and allows students to explore their interests within structured and individualized curriculum.” Similarly with AI, Empire Library employed task forces and an AI fellowship program for AI tool evaluation. Test runs and usability testing are used to review new AI offerings for elements including attribution, scalability, and accessibility. Pritting shares, “Accessibility specialists, faculty, instructional designers, and librarians reviewed what we needed to consider so that we could adopt AI tools to remain current while also communicating and engaging with the community before we adopt AI tools.” Pritting adds that “vendor partners in the library market can be great strategic partners and we should continue to work collaboratively with them on AI that meets user needs while keeping high standards in key areas such as privacy, bias, and environmental impact.”
Context is Key
“Empire Library’s approach to AI tool assessment is to keep the review appropriate to the context.” For instance, assessment can differ based on the focus of the AI tool—whether it’s intended for content generation, data analysis, or other purposes. Considering the level of change a tool will have on teaching, learning, and research is also a component of Empire’s evaluative strategy. When assessing the AI tool for its discovery platform, Empire Library “found that it didn’t change anything related to teaching and learning, but did improve our discovery system’s ability to return results from natural language,” Pritting says.
However, when applying that same evaluative process to the new AI layer in a digital primary source platform, Empire observed a potential shift in the way students would learn from it. “We thought this was worth creating a short task force comprised of faculty, librarians, and accessibility staff to review [the platform] and decide whether to disable the AI or not. That approach has served us well,” notes Pritting. “Using a framework to gauge how much change the AI will create to crucial components of teaching, learning, and research.”
Yet again, usability is key. “Any AI that makes our systems easier to use adds significant value as most of our students are working adults who don’t have time to get guidance on how to interact with systems that aren’t user friendly.”
As for collection development, Empire Library references the ICOLC Statement on AI licensing. Based on the AI tool’s intended use, the library decides the level of analysis needed and which stakeholders should be brought into conversation. In general, Empire Library intends to “participate in as many ‘just in time’ access models or tools as possible, especially if these help us reduce subscription costs.” One example is the library’s use of Libkey Discovery and Article Galaxy Scholar, “an unmediated article purchasing service that also offers some AI integrations with Scite.” Pritting underlines that, “With these two tools working together, our students, faculty, and staff get articles either emailed to them in minutes or access them right in their browser, and we don’t have to maintain a large subscription/complex license to get this access.” Going forward, Pritting predicts that “AI-informed access and purchasing tools” will continue to gain traction.
Making a Seat at the Table
AI discussions on campus tend to be at the administrative level. Empire Library’s work on AI evaluation, implementation, and licensing is slowly helping the library earn a seat at the table—and even influence institutional policy. Pritting shares that “the library has been invited to participate in developing AI Syllabus statements and discussions about how the University can use AI to make internal processes more efficient.” The library’s efforts are also aiding users by providing them with accessible and reliable resources that enhance research and learning.
Reading Room
A selection of multimedia content to fit your time commitment.
5 minutes
- Detecting AI-Generated Text: Things to Watch For. East Central College, 2025.
- 5 Telltale Signs That a Photo Is AI-generated, by Anna Louie Sussman. Kellogg Insight, Kellogg School of Management, Northwestern University, 2024.
- Evaluating Generative AI Resources: Separating the Tools from the Toys, by Rachel Hendrick. LibTech Insights, 2024.
10 minutes
- What Happens When People Don’t Understand How AI Works, by Tyler Austin Harper. The Atlantic, 2025.
- The World of AI, by Emily Udell. American Libraries, 2024.
- Artificial Intelligence: For Students: Methods for Evaluating AI. LibGuide by Iona University, 2025.
20+ minutes
- AI Tools and Resources: Evaluating the Reliability and Validity of AI Generated text and media. LibGuide by University of South Florida, 2025.
- Practical Applications of AI in Libraries, by Wyoming State Library. YouTube, 2024. Video below.
- AI for Legal Research: Use Cases, Benefits, Challenges & More, by Mitul Makadia. nasscom community, 2024.
- Testing AI Academic Search Engines (1): Defining the Tools, Aaron Tay’s Musings about Librarianship, 2025
- Testing AI Academic Search Engines – What to find out and how to test (2), Aaron Tay’s Musings about Librarianship, 2025
Quiz Time
Test your AI literacy knowledge!
Discussion Questions
AI gives us a lot to think about. Share your thoughts in the Leave a Reply section below!
- What metrics do you use to evaluate AI tools for your library?
- What do you advise students to think about when evaluating generative AI models?
Coming up
Next week, Evan Fruehauf, Assistant Librarian at the University of South Florida, writes about creating community around generative AI. And we learn about how Brown University’s Critical AI Learning Community is promoting thoughtful AI adoption across campus.
Take me to the Table of Contents.
Know someone who might be interested in this micro-course? Share the registration page with your colleagues! Make sure you’re registered to receive all eight weeks directly in your inbox.
Learn more about Clarivate and Choice’s LibTech Insights.
84 responses to “Week 5: Evaluating AI”
-
We do not AI policy in place yet, but I think when evaluating AI we should look at transparency, citations, pricing, data source, accessibility, ease of use, relevance and accuracy.
Students should be advised about data privacy, training, capability, accuracy – looking at facts and hallucination risks as well as bias and transparency. -
What metrics do you use to evaluate AI tools for your library?
I am not sure on the specifics of this just yet, unfortunately. I would think output is a big metric used for AI tool evaluation and potentially ease of use for students and faculty.What do you advise students to think about when evaluating generative AI models?
Based on past courses, I usually remind students, if they’re using genAI, to pay attention to the output and to double-check and triple check what ever sources are cited in the output. -
1.What metrics do you use to evaluate AI tools for your library?
AI metrics will include accuracy and precision.2. What do you advise students to think about when evaluating generative AI models?
Students must not rely on AI, and students are not responsible for accuracy -
It is a good idea to verify information that is received through any AI tool before using it as it sometimes gives wrong sources.
-
What metrics do you use to evaluate AI tools for your library?
I evaluate the accuracy of the output. I ask questions similar to these: Can the tool provide accurate, verifiable information consistently? Is it easy to use? Does it provide biased responses?
I will also consider cost, reputation, duplication of services etc.What do you advise students to think about when evaluating generative AI models?
Ensure that you verify all information independently. -
Learned a lot of technical foundation knowledge in this weeks coverage, specifically about tokens and processing capabilities.
-
What metrics do you use to evaluate AI tools for your library?
-key metrics include accuracy and precision.
What do you advise students to think about when evaluating generative AI models?
-students must not rely on the AI’s output. They should treat everything as draft until all are verified. -
We are not using metrics to evaluate AI tools that are embedded in the vendor products we subscribe to. But we did have an experience with a ProQuest product that provided an AI summary and instructors asked us to remove the feature from the interface, so we did. When working with students, I do think it is important to begin with “Is AI needed here?” rather than “How can AI help with this?” Students’ agency in the research process is on the line when they offload some of the cognition to AI, so it is an important part of their development of critical thinking.
-
What metrics do you use to evaluate AI tools for your library?
In my role as a reference and instruction librarian, I focus on the accuracy of the results. Is the tool response reflecting what the document is saying? The efficiency if the tool is also important. Does the tool include citation generator, download options, ease of use are also important. The databases are hard enough for students to grasp without having to also instruct them on how to use the tool. When the Proquest Research Assistant was made available most students I worked with didn’t want to use it because it was confusing.What do you advise students to think about when evaluating generative AI models? I usually will tell students to look for links to actual resources that the tool used and to verify the results as the tools often make errors. I remind them not to accept the results until they verify for accuracy.
-
For my library, we still don’t have any metrics yet, but personally in evaluating the AI we should stress in terms of security, reliability, accessibility, scalability and the usability. For students, they should have in mind that AI should be used for guiding them in generating the ideas for their research and they need to check carefully the resources provided whether its from authentic resources with the correct facts. They also need to learn the structure prompts to get better result.
-
Our institution has not yet developed formal metrics for evaluating AI tools. However, security remains a key consideration in determining which platforms are approved or denied. When guiding our students, we emphasize the importance of critical thinking and information evaluation. We remind them not to automatically trust AI-generated outputs, regardless of how confident they may appear, and to always verify information through methods such as lateral searching and seeking corroborating sources.
-
At my library we usually start with evaluating AI resources based on vendor especially since as a Academic Library we have many vendors who are already trying to implement AI into their products. We also evaluate heavily based on our own experiences with the AI, using it significantly before presenting to students as a resource.
We host workshops and classes that show students how not only different AI resources work and are different but also how the same AI can produce different results. We encourage them to think critically about what the AI is giving them and to ask questions of our staff when unsure. -
Of course the first factor to consider is the purpose for the acquisition. With the limited knowledge I have, the reputation of the company is important. As for the students, I advise that they carefully check the references and vaildity of the content to avoid hallucinations.
-
When evaluating AI tools for the library and the university as a whole, we consider factors such as pricing, reliability, accessibility, privacy, security, and legal issues. Another factor that was considered was whether the tool is already included or can be added to one of our current subscriptions.
We advise students that AI tools are not always accurate, so it is important to verify information using reputable sources. We also discuss potential biases, hallucinations, and privacy concerns that may arise. In addition, we emphasize academic integrity. Students should not use AI to complete their assignments, but they can use these tools ethically to support their learning. -
I try to learn about the tools by using them and assessing their ease of use and usefulness to the public, comparing the answers they provide with those of other tools: whether they are of high quality and whether they cite reliable bibliographic sources. But we never purchase any services.
-
Like any new information source check for quality (accuracy, clarity, credibility, source etc… ). Compare it with other reputable sources before using or including.
-
We do not have an AI policy in use yet, but the library is working towards it.
We advise the students to double check what AI is telling them. That there are a lot of hallucinations to work through. -
While we currently do not have an AI tool available for the organization, it is essential to emphasize the metrics for researchers who have personally used or purchased a subscription. With the best practices from this lesson, they will be equipped to address questions about AI when asked to contribute to the factors for consideration.
I want to highlight that it is crucial to develop a clear prompt to obtain accurate results for informed decision-making. -
Some of the metrics used to evaluate AI tools for my library are the sustainability, pricing, that is cost, reliability, user needs both for the students and professional staff, functionality, and the legal matters.
Students should be able to check out the usability and reliability of the model. How best the model will suit their needs comprehensively and qualitatively. -
I don’t have any metrics yet, we are working on an AI project to improve AI literacy and generate policies, for students and teachers.
-
¿Qué métricas utiliza para evaluar las herramientas de IA para su biblioteca?
La mayor parte hemos usado usamos formatos de IA, como ChatGPT. Para detectar los anexos de las tesis que tiene fotos y verificar su autenticidad. Seguridad del contenido: Identificar lenguaje inapropiado en respuestas generadas.¿Qué les aconseja a los estudiantes tener en cuenta al evaluar modelos de IA generativa?
• Verificar y analizar la información que sean lógicas y consistentes con lo que se consulta, debido que la IA pueden crear respuestas incorrectas o por lo que se aconseja consultar y verificar con otras fuentes confiables.
Protección de datos personales que tomen en cuenta que la información es sensible o personal ya que pueden almacenar o utilizar los datos. -
Our institution has not really established set metrics for evaluating AI at this point. Security, however, is a big factor for which platforms are approved or denied. Regarding what we currently tell our students, we focus primarily on critical thinking and information evaluation. Not trusting the output, even though it sounds very confident. To always double check the results with techniques such as lateral searching and looking for corroborating information.
-
I look at accuracy, usability, and sustainability first—how reliable are the outputs, how intuitive is the interface, and how sustainable is the vendor or platform? I also weigh privacy protections and whether the tool aligns with the library’s mission of equity and accessibility.
I encourage students to ask: Where does the information come from? And can I verify it? They should be aware of possible biases, hallucinations, and privacy risks, and remember that AI is a tool to support—not replace—critical thinking and traditional research methods. -
We don’t have an active AI statement at this point. We have used some ChatGPT, open AI formats. I found the article about detecting fake photos fascinating, especially when you look really hard, they are just odd looking. Does anyone call out fakes or see them on Social Media? That would be an interesting thread to read.
-
AI literacy is important to make informed decision on which AI tools are relevant, accurate and reliable. Information literacy for students is very important training students how to effectively use the tools and how these tools are supporting and advancing their learning experience
-
What metrics do you use to evaluate AI tools for your library?
I work in public services at my library, so I typically evaluate AI tools based on how helpful they are to students. I look at the user interface to determine ease of use and accessibility, and I look for accuracy in results. I usually test a new AI tool by asking it factual information and for research help to see if it produces hallucinations.What do you advise students to think about when evaluating generative AI models?
The #1 thing I tell students to think about when it comes to evaluating and using gen AI tools is intellectual property and copyright. I work mostly with multimedia, so most of the gen AI tools I help students with are image and video generators. I advise students to read the license and terms of agreement before they use a new tool so that they can understand a) where the generated content is coming from, and b) if and how their own work may be used for training or other tasks. -
What metrics do you use to evaluate AI tools for your library?
The criteria for evaluate AI: is transparency? provide trustworthy information? provide support to the institution?What do you advise students to think about when evaluating generative AI models?
appropriate use, academic integrity, critical thinking.
It’s so important to use a AI policy Manual. -
What metrics do you use to evaluate AI tools for your library?
• Is this AI Tool good for my academic community? It is correct for my researchers, professors and students?
• How does this application fit into an existing library collection?
• How does the application measure up against the developer’s or publisher’s claims of productivity and performance?What do you advise students to think about when evaluating generative AI models?
From all the readings in this week, the best that are close to what we ask our academic community are like the ones in Georgetown University …
When working with AI, keep in mind the following best practices for evaluation:
• Meticulously fact-check all of the information produced by generative AI, including verifying the source of all citations the AI uses to support its claims.
• Critically evaluate all AI output for any possible biases that can skew the presented information.
• Avoid asking the AI tools to produce a list of sources on a specific topic as such prompts may result in the tools fabricating false citations.
• Always remember that generative AI tools are not search engines–they simply use large amounts of data to generate responses constructed to “make sense” according to common cognitive paradigms. -
Whenever new tools are implemented, it is necessary to evaluate their function and impact in order to choose the best option and the best way to train students how to better use it.
-
At Chapman University, a multi-model AI tool, PantherAI, was just launched to students this fall semester I had an alternative to ChatGPT. PantherAI combines GPT-4 (via Azure), Claude (Anthropic), and Gemini (Google). The model is chosen based on the task, like advanced reasoning, coding, creative tasks, or visual reasoning. This tool had more security and privacy for University-affiliated users as chats are securely transmitted only to providers, with minimal retention on external servers.
Chapman University emphasizes AI literacy for students, and the library has I already thought of workshops and instruction to teach students critical and informed understanding of AI tools.
-
Metrics for Evaluating AI tools– accuracy, relevance, and reliability. Need to know if the tool provides correct and reliable results. Outputs….do they align with the query or task that was assigned. With the reliability of the AI tool being used are the results consistent over mulitple uses.
-
What metrics do you use to evaluate AI tools for your library?
Quality of output and, although it’s a double-edged sword, price is a consideration. Why? Affordability is important, but there are tradeoffs in what happens to the data. Sometimes the option to keep searches private is behind a paywall. So does that mean students with funds are the only ones who should learn how to use AI effectively? Should we shortchange students without a ton of money? No, instead we look at equity and share information so that users can make an educated decision. Example: Perplexity has a relatively robust free tier, and if users understand the implications of their use, they can still get a lot out of it.What do you advise students to think about when evaluating generative AI models?
1. What happens to the content you put in, whether it’s your questions or other people’s intellectual property?
2. Who is behind the model underlying the tool? Is it biased because it’s primarily content written by those with the privilege of getting their output published somewhere, or is it more representative of the population? How are the algorithms tuned?
3. What are the other potential impacts (learning, environmental, etc.) of your use of this tool? -
1 Output quality (accuracy, consistency, completeness of responses, level of bias detected…), usability and accessibility, security and privacy, scalability and technical sustainability (capacity for integration with library management systems, economic impact, ethical and environmental impact).
2 Verify reliability ( check the accuracy of the information by comparing it with academic sources databases, catalogs, repositories), review biases, privacy, traceability (properly cite AI responses as support, not as a primary source), technical limitations, responsible use (take advantage of it to support writing, generate ideas, or summarize). -
At this time, our institution just has an AI policy in place for Academic Integrity.
-
My advise depends upon the student level and the task for which they intend to employ AI. I suggest first that students determine their basic understanding of AI, then their desired output for the project. After creating a list of outcomes, I work with them on how they think AI can help them achieve these. Then we conduct a search of AI tools and capabilities, where I ask the student to decide whether the descriptions of the tools they find will actually produce the content they desire. Will the tool provide documentation? Is it trained in academic-quality data? Are there biases in the results? I help students understand that AI tools are trained on different bodies of data, making some much more useful than others. I explain what it means to use an OpenAI LLM. Finally, in working with students who are about to use AI, I do mention the environmental costs and ask the student whether they are willing to use other methods, even in part, that might reduce the energy footprint.
-
What metrics do you use to evaluate AI tools for your library?
What do you advise students to think about when evaluating generative AI models?Honestly, a lot of the time it comes down to the budget — is the AI tool open source and is the free subscription worth using (ChatGPT)? Is it included with our Microsoft email accounts (Copilot)… and then we can’t get away from making a Google search without Gemini popping in with an AI overview. It’s more a matter of what is available and already being used and how can we show students how to use these tools in a way that benefits them.
When evaluating generative AI models, I think students need to really understand foundational concepts of Information Literacy such as how to tell that a source is credible. They need to understand basic information types and view AI as a non-authoritative source like a wikipedia article.
-
I totally agree with you, Jill, when you mention recognizing what the students are already using. The focus should be on teaching them how to use the tools effectively and in ways that support their learning.
-
-
We don’t have a process yet, but so far we are evaluating AI enhancement of our discovery tool by simply testing it with sample searches and observing the differences. For students, we are emphasizing the need to verify and validate everything.
-
1). Since our campus libraries are in the early stages of considering the use of AI tools, we expect to start with traditional metrics (user needs and functionality, content and quality, cost, etc.) even as we consider issues such as ethical and legal factors, performance issues, technology sustainability, and vendor reputation etc. In addition, given the offer of additional vendor services such as AI research assistant we may have to engage our users to take part in trials, observe the users during those trials, and review the feedback from those trials to help understand our users’ needs.
2). Our IL sessions reveal low usage of AI tools. However, our instructors have advised the few student users of AI tools to verify the accuracy of the sources provided by those tools. In the coming semester we look forward to encouraging students to use more meticulous approaches for fact-checking outputs from AI tools such as lateral reading.
-
We use the same as what has been recommended in the readings. Does it perform the way we need it, how accurate are the results it produces, etc.
We tell the students to look out for the inaccuracies, to critically evaluate the product the AI produces, etc.
-
My advices about Generative IA includes not using the first answer obteined by the tool but reading carefully, on the other hand, I recomend to make different questions to the IA tool to discover if the tool give weird or biased answers.
-
1). Since our campus libraries are in the early stages of considering the use of AI tools, we expect to start with traditional metrics (user needs and functionality, content and quality, cost, etc.) even as we consider issues such as ethical and legal factors, performance issues, technology sustainability, and vendor reputation etc. In addition, given the offer of additional vendor services such AI research assistant we may have to engage our users to take part in trials, observe the users during those trials, and review the feedback from those trials to help understand our users’ needs.
2). Our IL sessions reveal low usage of AI tools. However, our instructors have advised the few student users of AI tools to verify the accuracy of the sources provided by those tools. In the coming semester we look forward to encouraging students to use more meticulous approaches for fact-checking outputs from AI tools such as lateral reading.
-
-
We don’t have a formal way to evaluate at my library. We use what was purchased for the entire college. It is still fairly new for us.
Students need to always make sure the information is accurate.
-
1. What metrics do you use to evaluate AI tools for your library?
We do not have a formal procedure that we follow to evaluate AI tools in our library. However, we do inform our users during information literacy classes to be cautious when using AI tools. The metrics I use to evaluate are accuracy and reliability (use other credible sources of information to check if the information provided there is accurate and reliable) Another aspect I use is the disclaimer (most AI tools does indicate that AI does make mistakes, hallucinate and provide false information. I also inform them to look for bias because AI tools can be bias and can also discriminate. I also inform them to think about ethics and about academic integrity (know what you are using the do not rely on it for almost everything especially your academic studies)2. What do you advise students to think about when evaluating generative AI models?
Even though we haven’t adopted a formal procedure in our institution – the following I would advise students to firstly think know that information AI is not 100% accurate and reliable. To safeguard themselves against this, they should compare with other credible and peer-reviews sources like journals, books etc. The second advice is that AI tools can hallucinate and provide information that is not credible and can also come up with references that are not there. The latter could harm their academic work if left unchecked, that is why it is important to check the limits and strengths of these AI tools. It is also important for students to know the AI model token are limited – this will help them to prompt the exact information to save time. -
We provide events and workshops, resources such as guides, FAQs and videos to support the evaluation of Gen AI usage.
-
I’m afraid our institution lacks a framework to evaluate AI, we only have AI policies. The information about AI we offer to students is only centered in its use but we don’t talk about how they must be evaluate.
-
What metrics do you use to evaluate AI tools for your library?
My institution doesn’t currently have a policy on assessing AI tools specifically, but we particularly emphasize transparency and privacy in terms of when we use AI tools.
What do you advise students to think about when evaluating generative AI models?
Consider the depth of analysis that a model provides. Know that each model works a little differently and may not have the same approach to solving a problem. Try using different prompt engineering frameworks (such as CLEAR) across models to see different responses. At this point in time I would try and show patrons how to use things like Perplexity to use research but would point out the different flaws and “hallucinations” that occur with any model and how to properly verify LLM outputs.
-
Queensland University of Technology Library has developed a framework for evaluating tools that we use as part of our workshops and self-paced modules for supporting HDR students to develop AI literacy skills. We are continuing to work on further improvements but it is open access and CC licensed and available at https://libguides.library.qut.edu.au/c.php?g=963920&p=7029563
Our rubric is adapted from Caico, M., Harris, L., O’Shea, S., & Mitchell, E. (2024). Evaluative information literacy rubric for AI tools. SUNY Oswego Faculty and Staff Scholarly Publications http://hdl.handle.net/20.500.12648/14992 (CC BY-NC) and University of Texas Libraries. (2024). Evaluating AI tools and output. Artificial Intelligence (AI) https://guides.lib.utexas.edu/c.php?g=1363366&p=10070755 (CC BY-NC).
We would appreciate any feedback. We would be keen to know if you re-use or adapt the QUT Library AI tools evaluation rubric.
-
Does the tool do what it says, how transparent is it on their about page or FAQs, is it open about how our data is being used, cost, purpose of tool, company behind the tool.
-
We have not decided or using metrics. When advising students evaluating generative AI models, it’s important to encourage them to think critically about reliability, bias , and appropriateness
-
Evaluating AI tools for the library:
To be honest, I am not sure what exact metrics my library uses to evaluate AI tools, since I haven’t been directly involved in those decisions. That said, I imagine that if my library were to adopt such tools, important factors would probably include how accurate and reliable the results are, whether the tool respects user privacy, and how easily it could be used by both staff and patrons. Cost and accessibility would likely also matter, since libraries work hard to balance budgets while keeping resources open and equitable.Advice for students evaluating generative AI models:
When it comes to students using generative AI, I would encourage them to look closely at how credible the information seems. Sometimes the answer AI gives sounds reasonable, but other times it can be so far off that it doesn’t make sense. In those cases, it’s important to double-check with reliable sources. I would also remind students to keep in mind our university’s academic integrity policies and to use AI in a way that supports—not replaces—their own work. Finally, I think it’s important to remember that AI doesn’t always capture every perspective fairly, so students should make sure they are considering a variety of voices and not letting the AI overrule more respected or authoritative sources. -
I don’t believe we have any specific tool recommendations quiet yet. We do have a guide on AI but I’m not sure how we came to those recommendations. When discussing with a student, I would tell them that they may use the tool for brainstorming and as a source but to still be critical of the results and give some examples of how the tool may be inaccurate in it’s recommendations.
-
Cross checking the results.
-
Aún no hemos adoptado una IA como recurso oficial en la institución.
-
What metrics do you use to evaluate AI tools for your library?
When it comes to evaluating AI tools I usually ask the following questions: Is this something I can do myself? Is the result of using this product something I will need to rigorously check to the degree it would be easier to do it myself the first time? What is the purported value add of this technology? Is that value add real? What is this tool claiming to do? Is that possible or does AI make the most sense for that job? Is the task this tool would be used for justifiable for the amount of environmental extraction it will be responsible for? Is the value this tool offers worth the social and labor costs that went into it? How will the use of this tool impact our relationship to information? To writers, artists and other creators? To each other?What do you advise students to think about when evaluating generative AI models?
I encourage students to really think about what they would be asking the model to do and what the cost is to them. As most students are not paying for higher scale models its important for them to ask why this tool is being provided to them for free, how it might differ from the more costly versions, and what they are bringing to the company if it is not payment. I also encourage students to think about what part of their thinking they are ceding to the system and if that seems worth it or if they think that aspect of thinking may be something they will need at a later date even for less mundane tasks. Lastly, I encourage students to really think about the environmental and social costs that come with using these tools over interacting with a classmate, professor, or other person whom they are or could be in community with. -
What metrics do you use to evaluate AI tools for your library?
The metrics are currently evolving at our institution. The librarians who handle our eResources are still shaping their approach to AI uses, focusing mostly on the AI attached to databases. They recently utilized inputs from faculty, asking the faculty to test out certain features to see if it met their needs or if it raised concerns. The librarians also considered ease of integration and use, accuracy of outputs, and whether the AI function was generative or agentive, which they prefer.What do you advise students to think about when evaluating generative AI models?
Students should understand that this can be a component of learning, but it should not replace learning. If it is spurring your ideas and helping to open avenues of thinking, then go for it. However, if the AI has done all of the thinking for you, and it feels more of a copy/paste, this direction is something to avoid because your institution likely has a policy against this type of usage, complete with consequences, and by not absorbing the information, you could likely have a less successful outcome in the class itself. -
What metrics do you use to evaluate AI tools for your library?
So far, we use AI tools from databases that have been recently integrating them (ProQuest). We immediately asked our vendors where does the AI tool take information from.What do you advise students to think about when evaluating generative AI models?
Observe the quality of the information and ALWAYS verify provided resources. If its based on hallucinated references then the quality of information is not reliable. -
1. We look at performance, cost, sustainability & environmental impact, legal implications, privacy standards & ethics.
2. We play test the “free” tools the students are using so we learn the pros and cons of each and so that we are coming from a place of familiarity during our instructional sessions. We tell the students to keep these key factors in mind when using Gen AI: always cross check for accuracy of information/listed sources, identify bias, ask “is the AI’s logic sound”, and finally does the AI tool use any emotional or manipulative language. How the tool is trained is another factor we often think about and mention but don’t have time to touch upon in great detail. -
1) What metrics do you use to evaluate AI tools for your library?
I am not sure exactly what metrics we use to evaluate AI tools, but I know my own library has workshops coming up that discuss how to use and evaluate tools. So maybe there will be something that comes up.
2) What do you advise students to think about when evaluating generative AI models?
I think first and foremost is for students to understand that these are tools and should have an idea of what they are going to be using them for. Certain tools are only useful in certain contexts, and students should be aware of that and know how to properly use them for that purpose. I think accuracy is one of the key factors when it comes to evaluation due to issues such as hallucinations or incorrect information. Double-checking the information is key. -
1. Unfortunately, as of the time I’m typing, we have no rubric for judging whether an AI tool is fit for our institution. Any AI in use is usually built into programs our students use already. However, as the landscape continues to change this will undoubtedly happen.
2. So far, we’ve had great success–mostly with our ESL and nontraditional students–in using ChatGPT in particular to help them create outlines (we’re attempting to convince them to use Claude instead, but ChatGPT is more recognizable). We also encourage them to try proving the AI tool wrong. We’re also working a section on proper use of AI into our library instruction for research techniques and for specific classes. So far there hasn’t been a test run, but we’re a small institution and change is slow here!
-
I would advise students when evaluating AI models to consider issues to dela with transparency, been able to confirm the accuracy, relevancy and ethical soundness of information, the efficiency of the tool in augmenting the research workflow rather than one which wastes the student’s time with biased results and hallucinations. They should also consider tools that enhance critical thinking skills rather than those that limits it. finally most important to the student is the tools privacy policy in order to safe guard their personal data.
-
I am yet to evaluate any AI tools in my library, but when that time comes armed with the knowledge already learnt through this course, the following will be my take. Just like many digital tools, I will evaluate AI tools by looking at , the user needs, functionality of the tool, quality of content, pricing, privacy considerations, ethical and legal factors, performance issues like accuracy, latency, scalability, and vendor reputation.
-
1. Students must look out for AI tools with the following features: article summary, exporting of tables and graphs, citation tool integration, what is the 3rd party data agreement, organisation of papers.
2. Quality of work which can be verified
-
When evaluating AI tools for our library, we focus on how well they align with academic integrity and instructional goals. We look for tools that support permitted uses of grammar correction, spelling assistance, and citation formatting. Tools that generate new text on behalf of students is more of a professor-by-professor basis, depending on the context of use. However, reliability, accuracy, and transparency are the most important criteria. The tool should clearly show suggested changes, give users control over acceptance, and consistently provide correct results. Above all, the AI tools should help students remain the true authors of their work.
When advising students on evaluating generative AI, I encourage them to start with purpose. They should ask whether this tool is helping to refine and polish my own writing, or merely creating content I can just copy-and-paste? Students need to think critically about whether a tool’s features stay within the boundaries of the code of conduct and academic integrity. I also remind them that their responsibility as scholars is to ensure accuracy, maintain ownership of their work, and use AI only as a support for clarity and presentation, not as a substitute for their original thinking.
-
When evaluating AI tools for our library, we focus on how well they align with academic integrity and instructional goals. We look for tools that support permitted uses such as grammar correction, spelling assistance, and citation formatting. Tools that are avoided are those that generate new text on behalf of students. Reliability, accuracy, and transparency are also paramount. The tool should clearly show suggested changes, give users control over acceptance, and consistently provide correct results. Above all, we ask whether the tool helps students remain the true authors of their own work.
When advising students on evaluating generative AI, I encourage them to start with purpose. They should ask whether this tool is helping me refine and polish my own writing, or merely create content for me that I can copy and paste. Using AI to generate new text is not permitted, so students need to think critically about whether a tool’s features stay within those boundaries of academic integrity. I also remind them that their responsibility as scholars is to ensure accuracy, maintain ownership of their work, and use AI only as a support for clarity and presentation, not as a substitute of original thought.
-
When evaluating AI tools for our library, we consider several factors, including costs, as they are a significant economic factor for a developing country’s budget. Then, we can look at accuracy and reliability, ease of use, integration with the systems the tool normally uses, privacy protections, ethical standards of the tool, and last but not least, good technical support and software updates.
I tell students that when they use generative AI, they have to always check the information generated to a trusted source. They should be aware of biases, and when using a tool, use one that can explain its process. They have to give credit to the AI when using it, they have to follow a set of academic honesty rules, and they have to know how much the tool costs. Generally, AI should enhance their thinking, not replace it.
-
The realm of scholarly communications, is a complex, evolving domain, beset with myriad challenges, particularly, the need to make capital of AI, Big Data and other emerging technological concepts (impacting appreciably on the understanding, authoring as well as ethical usage of publications in an increasingly techno-deterministic scholarly ecosystem), there is an increasing need for novel data-driven research evaluation metrics as per the recommendations of the Leiden Manifesto and the trepidation among academic circles in the wake of the publication of the 2022 OSTP Nelson Memo (mandating US federal grant agencies to draw up plans for making all federally funded research publications and data publicly available without any sort of embargo and delay by the end of 2025).
-
When evaluating AI tools for our library, we consider several metrics: usability, integration with existing systems, accuracy of results, bias and fairness, data privacy, scalability, cost, and environmental impact. We also look at how the tool improves workflows and supports both staff and patron needs.
For students evaluating generative AI models, I advise thinking critically about accuracy, reliability, and source transparency. They should consider potential biases, ethical implications, and the model’s limitations. It’s important to verify AI outputs, understand how the model was trained, and reflect on how its use affects their own learning and research integrity.
-
Metrics for Evaluating AI tools: 1) User/lecturer needs (this includes aspects such as functionality, quality, reliability, sustainability & performance)
2) Cost/Pricing
3) Legal considerations such as copyright and user privacy.
We only consider tools from reputable vendors that we have dealt with in the past.
We are still developing our response to AI and the information and library guides shared in this session will be very useful in helping to craft both guides and lessons. -
Currently, budgeting is the dominant evaluating factor–not only the cost of the tool’s subscription, but also how much it will impact our utility bills. In a year or two, budget will resume equal weight with privacy, relevant results and environmental impact. I advise high school students to actively ponder and discuss the AI tool’s output for relevancy, currency, authority (with easy-to-see citations). Then, I recommend that they never copy and paste the output but to take notes from it as well as the sources, in their own words.
-
We use a lot of the same metrics mentioned in other comments and in this article to assess which AI models we want to use, but ultimately, it’s up to whether or not it fits our patrons needs and is something they will actually use. If they are not excited about it and asking about it after a trial run, it’s probably not worth investing further time – and money – into providing long-term access to. Especially since we have institution-wide AI models available for people to use; whatever we add to the mix has to add value and not duplicate other efforts, and still remain affordable for us budget-wise.
-
Evaluating AI tools for libraries requires a comprehensive approach that considers not only functionality and workflow impact, but also content quality, privacy, sustainability, cost, scalability, and ethical risks. A balanced, evidence-based assessment ensures that the chosen tools align with academic needs and deliver meaningful value.
-
What metrics do you use to evaluate AI tools for your library?
The needs of our patrons (faculty and students mostly), budget constraints, the accuracy and reliability of the tool in question, ease of use, and whether it fills a need that we currently don’t have anything for or if it does a better job than another tool that we have.
What do you advise students to think about when evaluating generative AI models?
Again, does it fill a need the student has (here I talk with them about accuracy and reliability. If a tool doesn’t provide accurate information or resources then it isn’t fulfilling their needs as well as if it isn’t answering their question), cost (if any), currency of information/ resources, and any assignment or policy restrictions as to the use of AI.
-
To take into account the purpose of the tool, accuracy, fairness, transparency and all the ethical implications. how easy is it to use, privacy policies, integration and reproducibility capabilities. Licensing and Agreement and the pricing plays an important role.
Students should always do cross-checks if references are proved by the tool, critical evaluate and look out for biases. -
Alway question the accuracy in terms of real references , reliability and currency with world/topical events, privacy concerns and ‘access & equity’ issues . A student than can pay for a sub to an AI tool may get more ‘bells & whistles’ and less hallucinations than those that rely on free access. So, perhaps those that can pay get a better level of AI service. It’s important to have ‘play time’ and evaluate a variety of AI applications – take them on a test drive – include the Toy types as well as the Tool types.
-
The advice I give students on evaluating AI is similar to the advice I give on evaluating any information tool: evaluate content in context, consider the source, did this information resource or tool make you more or less efficient as a research?
I would rather students be thoughtful and critical of AI than “good” at it. -
When my turn comes to evaluate, I’ll focus on Ethics, Accuracy, Latency, and Scalability. And that’s what I’ll recommend.
-
In evaluating AI tools for our library, I focus on a combination of usability, accuracy, and ethical considerations. I assess how a tool affects our workflows, its reliability in producing verifiable information, potential biases, scalability, accessibility, and its environmental footprint. Cost and technical support are also important, especially given limited IT infrastructure.
When advising students, I emphasize that AI outputs should never be taken at face value. They should consider the tool’s accuracy, potential bias, transparency of sources, and ethical implications, and cross-check results with authoritative databases or scholarly resources. Encouraging iterative use, reflection, and critical evaluation helps students build both AI literacy and responsible research habits.
-
What metrics do you use to evaluate AI tools for your library?
Privacy has been a big concern, and of course accuracy. I am aware that biases exist, but was unsure really how to check for that. Library-specific vendors would seemingly be a good place to start. -
As we don’t have an AI policy yet, we don’t yet have this codified. However, the literature presented in this section of the course will be very helpful in determining many of those guidelines.
-
Performance, costing, vendors reputation, reliability of report, legal issues should be considered carefully to evaluate AI tools. I would like to suggest students to assess the easy usage and the accuracy of delivery of the AI tools to evaluate the tools
-
The accuracy of the results is extremely important. This includes both the content and articles. It is important to check the articles to see that they exist and if they are reputable if you are looking for articles. This is also true of books. Sometimes the books don’t exist.
Accuracy includes images. Strange images can be generated with AI.
Also, the amount of results is important. There should be a decent amount to choose from. Sometimes, there are not a lot of results.
In addition, currency is important. Many artificial intelligences will not pull from the web or pull from recent training data. Having recent training data is helpful.
There are also potential problems with bias. Some of the articles may contain biased information.
In addition, protocols like RAG or MCP as well as the ontological structure being used are important for accuracy.
-
What do you advise students to think about when evaluating generative AI models?
When evaluating generative AI models it is most important that students use critical thinking skills that encompass fact-checking to ensure information is from reliable sources and that information is referenced to real sources. They should also be aware of potential bias that may exist in outputs. Another important consideration is potential limitations such as misinformation or hallucations that can and do occur. Being intentional which how and why it was used and then evaluating its effectiveness should remain central to any interaction with AI models
-
What metrics do you use to evaluate AI tools for your library?
Output accuracy is a crucial metric when evaluating AI tools. If the tool consistently produces accurate results with attribution of sources, then it is possibly a useful tool. Verifying the accuracy is still the responsibility of the AI tool user.
What do you advise students to think about when evaluating generative AI models?
As I mentioned above, students are always responsible for the accuracy of the results provided by an AI tool. Therefore, verifying the results is critical.
Take me to the Table of Contents.
Leave a Reply