Austin, TX

Ian

Advanced
@ian-nai

Librarian at UT Austin.

In-Browser-OCR. A web app for performing OCR on images within your browser.

57

PDF-Scraper. Python scripts to extract text from PDFs, save it as a text file, export a list of words and their frequencies to a CSV file for further analysis, extract dates from the text, and graph the text's parts of speech.

35

PyGallica. A Python wrapper for the National Library of France's Gallica API.

22

rust_readability. A package to assess the complexity of texts using a variety of readability formulas.

9

Simple-Web-Archiver. A simple, GUI web archiving tool in Python.

8

REEES-Digital-Humanities. This repository aims to host content and tools relevant to the Digital Humanities, especially as they may be used to further scholarship in area studies and distinctive collections within libraries and archives. This aims to address both institutional and public needs, breaking down barriers between the academic and the public by providing open source resources for free use by all. The repository is currently divided into folders focusing on digital humanities tools and open access resources. More content will hopefully be added in the future, and I am open to any recommendations for additions.

8

Europeana-Full-Text-in-Python. Various Python scripts to assist with searching and downloading full text records via the Europeana APIs.

7

rust_lemmatizer. A lemmatizing package written in Rust.

6

french-syllabification. A rule-based algorithm to split French words into syllables.

5

Simple-Sentiment-Analysis. Simple scripts for performing sentiment analysis in Python.

4

rust_stylometry. A package to perform stylometry operations, written in Rust.

4

domovyk. A Python package to transliterate various Cyrillic alphabets to and from the Latin alphabet using the American Library Association-Library of Congress Romanization tables.

2

Non-English-NLP-Tutorial. Tutorials on processing and analyzing text using NLP methods.

2

Rozha. A Python package for performing multilingual NLP.

2

WebScraper-Cleaner. A Python script that lists all pages linked from a web address, scrapes an address specified by the user, exports that page as text, then removes script tags from the page and saves the cleaned product as another text file. BeautifulSoup and lxml are used to parse and clean the HTML and text.

2

dewey. An esolang based around the structure of the Dewey Decimal System.

1

sukhareva. A website highlighting the contributions of Soviet psychiatrist G.E. Sukhareva to early autism research.

1

wordfinder. A word finding package that generates words containing specified letters. Comes prepackaged with an English dictionary, but can also be used with custom dicitonary files and arrays of words.

1

raku_lemmatize. A Raku module to lemmatize strings and lists.

1

SciPy2023-Poster. My virtual poster presentation for SciPy 2023.

1

viterbi_pos_tagger. Rust

1

wordgoblin. A simple, lightweight word finder package that returns words containing letters specified by the user.

1

PyGematria. A Python package to facilitate numerological analysis on historical and contemporary texts.

1

NLP-on-Khlebnikov. NLP on works by Velimir Khlebnikov.

1

tx_data_repo_tools. Code to help with various tasks related to the Texas Data Repository.

1

DPLA-Map-Search. Search the DPLA API by setting a marker on a map.

1

RSS-Viewer. Simple RSS parsing and viewing in JavaScript.

1

Negritude-UT-Austin. A site containing resources and information regarding the Négritude movement.

1

Language-Splitter. Python scripts to detect languages in a text and split the text into its component languages.

1

OA-Journal-Generator. A web interface to generate the title, URL, and JSON metadata for a random open access journal given user-provided keywords.

1

omeka_saver. Save Omeka exhibits in Python.

1

DOAJ-in-Python. Python wrapper for searching/downloading metadata through the Directory of Open Access Journals API.

1

NLP-on-Daumal. NLP on novels by René Daumal.

1

Analyzing-Peguy. NLP on works by Peguy.

1

Twitter-Friends-Scraper. Scrapes a list of a Twitter user's friends (accounts they follow) and exports them to a csv.

1

krupskaya. An online resource for information on Nadezhda Krupskaya.

1

Musical-Chess. Music generated by moves on a chessboard.

1

HathiTrust-API-Searcher. Search the HathiTrust API for bibliographic records using OCLC, ISBN, or HathiTrust Volume Identifier numbers.

1

PartofSpeech_Grapher. A Python script to graph the distribution of parts of speech in a given text.

1

Date-Scraper. A simple Python script to scrape dates from a given text and save them to a .txt file.

1

Russian-Ebooks. A collection of full-text Russian ebooks, sorted by author, for reading, text analysis, and other scholarly and digital humanities work.

1
41
Apply