This is your work, valued
Computational literary analyst. I write code that helps us to understand novels and poems.
text-matcher. A simple text reuse detection CLI tool.
140chapterize. A simple tool for splitting up an ebook into its chapters. Works well with Project Gutenberg texts. May also be used to clean up books for computational text analysis.
120course-computational-literary-analysis. Course materials for Introduction to Computational Literary Analysis, taught at UC Berkeley in Summer 2018, 2019, and 2020, at Columbia University in Fall 2020, and again at UC Berkeley in Summer 2021 and 2022.
93workshop-text-analysis-spacy. Materials for the workshop Advanced Text Analysis with SpaCy and Scikit-Learn, given at NYU during NYCDH Week 2017, at PyData NYC in Nov. 2017, and at Columbia University in 2018 and 2019.
82corpus-db. A textual corpus database for the digital humanities.
64dotfiles. My personal dotfiles, using Nix Flakes to configure my system(s).
40macro-etym. A tool for analyzing the word histories of a text.
37gitenberg-experiments. Scripts for scraping metadata from Project Gutenberg books, via GITenberg.
20corpus-list. A structured list of text corpora, created for use with a corpus downloader.
13late-style-PCA. An attempt to experimentally test Edward Said's claims about late style using computational text analysis and principal component analysis.
11allusion-detection. Computational intertextuality detection in Python. Fuzzy string matching, approximate string matching.
11cenlab. A corpus of English-language novels combining the ~250 novels of the Corpus of English Novels with the Txtlab corpus of English novels.
9course-computational-literary-analysis-readings. Syllabus and course readings for Introduction to Computational Literary Analysis, a course taught at UC-Berkeley in Summer 2018, 2019, and 2020, and at Columbia University in Fall 2020.
9md2mla. A script and accompanying templates to make an MLA-style paper from a markdown file. Requires Pandoc and LaTeX (xetex)..
8workshop-word-embeddings. Materials for a workshop in word embeddings, for NYC-DH Week, February 2019
7book-computational-literary-analysis. A textbook for the course, Introduction to Computational Literary Analysis. WIP
7workshop-dataviz-2017. An Introduction to Text Analysis and Visualization, Art of Data Visualization Week, April 2017, Columbia University
7milton-analysis. Text analysis of Paradise Lost and other poems by John Milton.
7dissertation. A dissertation in computational literary analysis, called "The Eye of Modernism: Visual Imaginations of British literature, 1880-1930"
7jonreeve.com. My personal website, jonreeve.com, written in Haskell, using Ema.
6template-research-paper. A template for a research paper, which compiles to many file formats.
5shakespeare-dialog-extractor. An application to extract dialog from Shakespeare plays, as encoded into TEI by the Folger Library.
5docmap. A project for creating new themes and customization functionality for the Omeka content management system.
4free-indirect-discourse-model. Modeling free indirect discourse in literature, using AI.
4conference-joyce-digital. Website and materials for the conference Joyce in the Digital Age, held at Columbia University on October 1st, 2017.
4course-cic-compling. Course materials for the course Computing in Context section in Computational Linguistics. Dept. of Computer Science, Columbia University, Fall 2021. Work-in-progress.
4template-dissertation. A template for a modern, best-practices dissertation.
4plato-analysis. Analyses of Platonic dialogues, including a Socratic dialogue generator.
3course-word-embeddings. Course materials for "Meaningful Text Analysis with Word Embeddings," taught at the Digital Humanities Summer Institute, June 2021.
3dissertation-prospectus. My ever-protean dissertation prospectus.
2htrc-experiments. Text analysis experiments with Hathi Trust Research Center literary datasets.
2template-course-website. A website for a university course. Semantic by default.
2occupations-experiment. Experiments in quantifying occupations as they're represented in fiction.
2corpus-joyce-portrait-TEI. The Open Scholarly Edition of James Joyce's A Portrait of the Artist as a Young Man
2sops. Research materials (literature review, bibliography) for the project A Safer Online Public Square
2text-to-time-series. Experiments in text analysis, generating time series from texts.
2html2tei. A tool to extract structured data from novels (starting with Project Gutenberg HTML files)
2david-copperfield. An annotated edition of David Copperfield
1pg-srp. Stable Random Projections (SRP) of Project Gutenberg texts, for similarity tests
1hs-tei-transform. Experiments in transforming TEI XML, using Haskell
1course-nyu-pit. Course materials for the New York University Institute in Public Interest Technology (NYU-PIT)
1chaucer-macro-etym. Macro-etymological analyses of the Canterbury Tales.
1plugin-CollectionTree. Gives administrators the ability to create a hierarchical tree of their collections.
1course-multilingual-technologies. Course website for Multilingual Technologies and Language Diversity, taught at Columbia University by Prof. Smaranda Muresan and Dr. Isabelle Zaugg
1plugin-ExhibitBuilder. Allows users to develop online interpretive exhibits that combine items in the Omeka archive with narrative text. The plugin provides pre-built themes and layouts, and a WYSIWYG visual editor, to build complex pages.
1dataviz-workshop. Materials for a workshop in text analysis and visualization, originally given at Columbia University in April 2016.
1persistent-homology. Experiments with NLP and persistent homology.
1course-university-writing. Draft materials for the course "University Writing with Readings in the Data Sciences," taught at Columbia University in the fall of 2017. Students, please refer to CourseWorks instead of this repository.
1org-autolinks-mode. An emacs minor mode for automatically linking to org files, after typing the name of the file.
1data-ethics-literature-review. An automated survey of literature and curricula surrounding ethics in data science. WIP.
1