Beyond the Job Title: Digital Initiatives Librarian

A library tech job interview with Eric Lease Morgan

Job banner reading Digital Initiatives Librarian

Core to LibTech Insights’s mission is demystifying the broad and dynamic field of tech librarianship in higher ed. In this series, we interview a librarian every month to learn a little more about their position. As library tech jobs proliferate, they sometimes come with unfamiliar, jargony, or intimidating titles. We want to go beyond the title and look at the responsibilities, skills, and joys that make up the job. We hope this series will increase your knowledge of library tech jobs and skills and offer you greater insight into the working lives of your colleagues.

For this installment, we spoke to Eric Lease Morgan to learn more about his job as a Digital Initiatives Librarian. Check out our archive of job profiles. 💫


What is your job title and responsibilities?

I believe my official title is Digital Initiatives Librarian, but I just call myself a librarian. I have three responsibilities:

1. Teach/facilitate workshops on the topics of text mining and natural language processing. Over the past number of years, I have taught/facilitated about three one-hour workshops per week during the academic year. These workshops include introduction to natural language processing (NLP), Python and NLP, concordancing, topic modeling, how to use the Distant Reader, how to make a book, and how to write in a book. By now, I have easily instructed hundreds of students per year.

2. Work collaboratively on projects with undergraduates, graduate students, and faculty. All these projects include components of text mining. Example research questions have included:

  • Given a set of 2,400 perfect citations, which ones should I actually read for my thesis?
  • How do students discern their majors?
  • What human values exist in Bitcoin discussions? What human values exist in world constitutions? How do those human values compare and contrast?
  • To what degree do the ebb and flow of fentanyl deaths affect negative perceptions of China?
  • How has science changed over space and time?

In all of these cases, I first amass and curate large collections of text, and these things are really datasets. (Think “collections as data.”) I then model—analyze—the datasets to address the research questions. For example, I have harvested 750 books from HathiTrust, and each book was 750 pages long. I have used the Nexis Uni API to download 48,000 newspaper articles. I have worked with a graduate student to curate a collection of 600 Victorian novels. In these ways, I practice every aspect of librarianship: collections, acquisitions, cataloging, preservation, and dissemination. In many of these cases, the results of the research are journal articles, and I appear as a coauthor.

3. Investigate how computer technology can be exploited to improve the processes of librarianship. I have been doing this for the whole of my 40-year career. Examples have included: information retrieval, personalization, usability, and artificial intelligence (AI). These things have been manifested as the Mr. Serials Process, the Alex Catalogue of Electronic texts, and MyLibrary. I wrote my first AI program in 1992, but at that time, AI systems were called “expert systems.” Recently, I have gotten a proof-of-concept grant from Amazon to explore how generative AI can be used in libraries. During the pandemic, I was awarded close to $.75 million dollars of services from Microsoft to collect, curate, and analyze 1 million scholarly journal articles on the topic of COVID-19.


🔥 Stay up-to-date with LibTech Insights by signing up for our free newsletter. Just one weekly email with our new blog posts, top tech news stories, and other bonus content. Check out some posts from our archive:


Can you give an example of a project you’ve worked on as Digital Initiatives Librarian?

For the past five or six years, I have been developing an ecosystem called the Distant Reader. Given an almost arbitrary amount of content of almost any type, the Reader creates datasets that can be analyzed in a myriad of ways. To demonstrate the Reader’s functionality, I have created a collection of 3,000 such data sets—affectionately called “study carrels.”  Moreover, to make it easier for people to create study carrels, I have created a collection of .7 million items with two different interfaces: a traditional catalog and a less traditional but more functional index.  Academics are expected to read a lot. The Reader facilitates the process; the Reader makes it easy to get one’s head around dozens of books or hundreds of articles. Believe it or not, the hard part is actually getting the content, not the analysis.

In a similar vein, I have helped other librarians use computers to do library better; more or less, I started and fostered the Code4Lib community. It began as a mailing list in 2004, and it has matured to include an annual conference, a refereed journal, and a number of regional Code4Lib communities. Currently, the mailing has about 3,900 subscribers. Larger than LITA used to be?

Can you describe your career trajectory leading up to this position?

In 1984, I was on the lending side of interlibrary loan at Drexel University, and I wanted to be a reference librarian. My boss said, “You will have to write an annual report; you will have to count and tabulate all of those little pieces of paper.” Well, I hated counting and tabulating little pieces of paper, so I wrote a program that created my annual report daily. It was then that I learned how computers could be exploited in libraries. It changed my trajectory. I now wanted to be a “systems librarian.” Later, I went on to be a medical librarian, and I got grants from Apple Computer and the National Library of Medicine. That was when I wrote my first AI program. I outgrew that job, and then I worked at the NC State Libraries, where I was one of the first 10,000 people in the world to create a website. Really. I outgrew that job too, became webmaster here at the University of Notre Dame, and after 20 years, I now work in a digital scholarship center doing the things outlined above.

As a humanist with a liberal arts education, I have always been interested in “great ideas”: truth, beauty, honor, justice, love, art, science, philosophy, religion, government, history, etc. I became a librarian as a way to be immersed in a profession where these sorts of ideas can be actively investigated and explored. I now have millions of items in my collections on these great ideas, and these items are measured in the multibillions of words. Moreover, I have both access to large computers—computers the size of Walmart—as well as the skills to use them efficiently. As I head toward retirement, I see my investigations continuing. The pursuit of truth and beauty never ends.

As an educator and a veteran in the field, what’s something you’d tell library school students who are interested in pursuing tech-focused librarian positions?

There are a few things, not listed in priority order:

  • In a computerized environment, the organization of data and information is made real via serialized data structures, and success in this regard is determined by two things: (1) knowing the strengths and weaknesses of various data structures, and (2) knowing how to transform data and information from one structure into another. One of the more common data structures in Library Land is MARC. It is great for organizing bibliographic information, like books and journals. But it is really, really poor for things like music and images. Other data structures are simple lists: apples, cherries, pears, bananas. There are more complicated lists: fruits (apple, pears, cherries, bananas), vegetables (lettuce, celery, onions, potatoes), meats (beef, chicken, fish), dairy (milk, cheese, butter). The former list is “flat,” and the second is hierarchal. Unfortunately, most of us librarians try to put these sorts of things into a spreadsheet sort of thing. [Eric slaps his forehead.] I could go on and on, but students ought to understand the differences between “flat” lists, MARC, matrices (“spreadsheets”), XML, JSON, and relational databases. If they can draw an “entity relationship diagram,” then they will be well along their way to understanding data structures.

  • Any computer solution is mostly about people. Yes, the use of this data structure over that data structure may make a lot of sense, but getting everybody involved to understand that is the hardest part. It takes a long time and sincere efforts on everybody’s part. The solution is effective communication. Go the extra mile to explain this, that, and the other thing. Do your best to stay away from jargon. Be patient.

  • In the current environment, “find” is not the problem to solve. You, I, and everybody else can find bunches o’ stuff. Such is easy. Instead, the problem to solve is threefold: (1) getting the things found from search, (2) prioritizing the acquired items, and (3) using and understanding—reading—the acquired items. Suppose you do a search for “climate change” in any given bibliographic database. Suppose you narrow your query this way, that way, and the other way. You use all the appropriate limiters and fields. You use controlled vocabularies. In the end, I’d guess, you still have dozens, if not hundreds, of resulting articles. Getting the items is problematic. How do you do it? Click. Save. Click. Save. Click. Save. You will get tired after just a few items. Now, suppose there is some sort of tool allowing you to download all the articles. How are you going to read them? The traditional reading process will not work. It is not scalable. We—as a profession—need to advance the science here. In the words of Ranganathan, we need to save the time of the reader, and computer technology is part of the answer.

  • Don’t conflate the types of information with the formats of information. The former includes books, journal articles, datasets, images, speeches, sounds, videos, paintings, etc. The later are things like PDF, HTML, plain text, JPEG, comma-separated value files, XML, JSON, etc. For example, an image can be manifested in any number of formats: GIF, PNG, JPEG, TIFF, etc. A dataset can be manifested in a whole bunch o’ ways. Then ask yourself, “Can a book really be manifested as a PDF file?” I say not, unless the PDF file is subsequently printed and bound.

You’ve had such a rich career, but what would you say has been the most gratifying part of it?

  • It was gratifying to see some of my applications run for decades sans any software maintenance

  • It was gratifying to have the opportunity to travel all over the United States and Europe giving talks and facilitating workshops

  • It was gratifying to have had the opportunity to work for so many great institutions

  • It was gratifying to collaboratively work with undergraduate and graduate students on so many interesting projects

  • Finally, it is gratifying to know I am well-versed in just about every aspect of librarianship; I can do just about everything from curating collections to cataloging, to archiving, to bookbinding, to providing public service to people in-person as well as worldwide