Five Upcoming Academic Books on Data Work, Big Tech, and Misinformation
Our most hotly anticipated books for fall!
Posted on in Blog Posts
Posted on June 24, 2024 in Blog Posts
Authors:
David Erlandson
In April 2020, the Program for Cooperative Cataloging began a pilot project to explore Wikidata to support of one of its strategic goals at the time: to “accelerate the movement towards ubiquitous identifier creation and identity management at the network level.” Wikidata was created to be a structured data commons—a place where humans could create and interact with machine readable linked data. While they primarily support other Wikimedia projects, the data are freely available to anyone. The PCC pilot project recognized that Wikidata had naturally grown into a popular identifier hub. Scroll to the end of any Wikidata entry and you’ll find a list of unique identifiers for how the concept is represented in other ontologies, making it useful for accurate disambiguation. Perhaps the big motivating point participants wanted to explore was whether Wikidata could be an authority-lite system. This would allow librarians to mint or reference discrete URIs without having to go through a formal NACO/SACO pipeline.
Seventy-seven institutions participated in this project, which involved learning linked data concepts, making edits to Wikidata, and exploring SPARQL queries. Each participating institution was encouraged to have its own practical component. These components could vary quite a bit. For example, one project at Stanford focused on (among other things) a digitized movie poster collection to make their system identifiers available for query by other systems. Others wanted to build up representation of their university’s names and locations. As I would find out later, this was one of the wisest first steps you could take in Wikidata. If you want to describe people as they relate to specific departments in your institution, you first need to have an accurate description of your institution’s structure (university, colleges, departments, programs, etc.).
In the case of Rice University’s Fondren Library, we have a running CSV file of historic names. The entries listed relate to Rice or local Houston history. Each line in the document represents something influential to the university—astronauts who attended Rice, influential department chairs, particularly famous alums, donors, etc. Prior cataloging staff created this thesaurus, which was usually consulted when historical photographs were digitized. This allowed the metadata in our repository to be kept in check, as there weren’t great system controls at the time. There are many downsides to keeping a standalone document like this. Primarily issues relate to access and versioning; you may soon find multiple departments have their own unique and edited copy.
It seemed that publishing entities from this thesaurus on Wikidata would resolve some practical issues while also allowing the library to embrace linked data. We could then harvest the URIs, which we made into contributions, and link them in our institutional repository.
If you’re looking to get started on Wikidata for your library, review this list of project ideas for inspiration.
🌟 Subscribe to the LibTech Insights newsletter for weekly roundups and bonus content, including:
Anyone getting started on a Wikidata project will need to take stock of what already has a representation on the platform and what needs to be created. To do this, one potent tool is OpenRefine, particularly its reconciliation service.
Suppose you have spreadsheet of names. OpenRefine can run against each name and provide options for how your data align (or don’t) with Wikidata. (Side note: reconciliation can be done against many data sources, such as VIAF, the Homosaurus, or the Library of Congress). You may use this to align your data. The example below shows how someone might review and select an option for an entry “Paris”. It could align with the capitol of France, a city in Arkansas, or town in New York or even in Illinois.

Once you know what already has an entry in Wikidata and what does not, you can begin cleaning or enhancing your data to ensure they’re ready for publishing on Wikidata. You may want to contribute additional properties to existing entities; alternatively, you may be creating completely new entities. This can also be done through OpenRefine, reducing how many systems you’ll need to interact with along the way.
You may want to set up certain columns that help describe the entity. I found it helpful to align additional column headers as property names. I could then start to build up linked data triple statements (i.e. “Kinder Institute for Urban Research” has parent organization “Rice University”). The image below demonstrates how data might look if you wanted to build up entities describing a college and its constituent departments.

Some columns will have unique entries in each, such as the departments specializing in a specific field like English or art history. Some columns will have repetitive elements because each can be described as an academic department in Houston, USA.
If you are embarking on a journey to publish to Wikidata, it’s only fair at this point to note that cleaning and reconciling your data can be difficult and time-consuming work. Plenty of options exist to handle writing your records in batch. However, the act of reconciling takes a fair amount of research. Deciding which entities are aligned with another system is not always easy due to the classic problems that authorities are designed to alleviate. That is, certain entities may look the same due to one-to-one matches by name. John Smith, the South African rower, is different from John Smith, the Canadian rower, who may be different from the John Smith in your dataset. The heavy lift is aligning your entity or differentiating it from any other existing one.
The exciting part of this program was being able to easily publish data in a way that other users found interesting and engaging. Publishing to Wikidata allows other users to point to and build on that knowledge. For example, a fellow by the name of Brian O’Brien worked at Rice University. But he also worked at the University of Iowa and the University of Western Australia, making contributions to physics everywhere he went. This builds an important picture of his life and allows each institution to recognize a connection. You might also be able to build up a list of awards this person has won or books they have published.
Beyond that, exciting things started to develop when you allow others to contribute to your knowledge. When describing Rice University on Wikidata, I felt it was important to mention some milestones in its history, both good and bad. One pride point is that our university has had mixed-sex education from its inception in 1912. On the flip side, we were a little late to racial integration—1963. I created the concept of school integration on Wikidata to describe this aspect of desegregation. What has been a pleasant surprise is that other users have found this helpful, applying it as a topical description here and there. Still other users have been contributing language labels. These are languages I do not know, such as Latvian and Slovenian, but anyone can use those terms now.
The PCC pilot had a dashboard to track statistics related to the program. Contributions totaled approximately 195,000 articles edited and 55,000 new entries created. And people were looking! Entries created or modified by the pilot participants were viewed more than 4.24 million times. Standout programs are from the University of Washington and the Frick Art Reference Library, having a combined contribution of just over 80,000 references added to Wikidata, nearing 2.3 million views.

As for Fondren Library’s program, we started small and pivoted a few times. I had wanted to begin with descriptions of people, but I found it necessary to build out an organizational structure for the university first. Through this journey, I coordinated with an MLIS practicum student (who, I’m also happy to report, is now managing metadata at the University of Houston). What I found was that it was relatively easy to learn and to teach interacting with Wikidata. We had a good cadence of weekly contributions. The actual number was relatively small compared to the standout programs, at 671 references added. But the impact was large at just over 30,000 views. My takeaway lesson from this was that small, well-planned contributions can make a lasting impact.
The LD4 Community has been promoting good linked data practices in libraries and archives for quite a while. Because of its work, an increasing amount of production tools are gathering linked data from the Library of Congress, enhancing them with Wikidata attributes, and making them practical for library patrons through things like person entity pages. This allows patrons to seek and view people relevant to their research in a more natural manner. It also allows a librarian who contributes to Wikidata to have a practical impact on their own library’s discovery layer.
I recommend making contributions to Wikidata to learn and deploy linked data. While there is a learning curve to any new technology, Wikidata has been designed to make things relatively easy. There is some research work involved, though it did not unduly dominate my time. I was very happy with the reporting results from our Wikidata efforts. There are many times where metadata projects happen but measuring their impact is difficult. Wikidata will notify you when something you’ve created gets linked elsewhere by the community, giving you a pleasant surprise every now and again after the project is complete.
🔥 Sign up for LibTech Insights (LTI) new post notifications and updates.
✍️ Interested in contributing to LTI? Send an email to Daniel P. at Choice with your topic idea.
Our most hotly anticipated books for fall!
Posted on in Blog Posts
What if work could feel remoralizing?
Posted on in Blog Posts
Insights and best practices for teaching AI literacy to history students
Posted on in Blog Posts
What our micro-course participants had to say about AI in libraries
Posted on in Blog Posts