Research engineer: Datasets, information retrieval, representation learning, LLMs, scientific & legal document processing
awesome-document-similarity. A curated list of resources on document similarity measures (papers, tutorials, code, ...)
256pytorch-bert-document-classification. Enriching BERT with Knowledge Graph Embedding for Document Classification (PyTorch)
161scincl. Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings (EMNLP 2022 paper)
79llm-datasets. A collection of datasets for language model pretraining including scripts for downloading, preprocesssing, and sampling.
66aspect-document-similarity. Implementation, trained models and result data for the paper "Aspect-based Document Similarity for Research Papers" #COLING2020
63legal-document-similarity. Legal document similarity - Code, data, and models for the ICAIL 2021 paper "Evaluating Document Representations for Content-based Legal Literature Recommendations"
32semantic-document-relations. Implementation, trained models and result data for the paper "Pairwise Multi-Class Document Classification for Semantic Relations between Wikipedia Articles"
31clp-transfer. Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning
30awesome-prompt-optimization. A curated collection of resources for prompt engineering, optimization, and automatic prompt generation across text, image, video, and multimodal AI systems.
18awesome-anonymization-for-llms. A collection of resources for PII detection, anonymization, privacy-preserving techniques, and GDPR compliance in Large Language Model (LLM) or AI applications.
18german-language-models. A collection of German GPT language models
11aspect-document-embeddings. Code, dataset & models for the paper Specialized Document Embeddings for Aspect-based Similarity of Research Papers (#JCDL2022)
11awesome-contrastive-learning-for-nlp. A collection of papers about contrastive learning for natural language processing.
7wikipedia-article-recommendations. Survey data and Python code for the ICADL 2021 paper "A Qualitative Evaluation of User Preference for Link-based vs. Text-based Recommendations of Wikipedia Articles"
5getting-started. Dockerfile
4covid-vaccination-appointment. Python
3Leaflet.Sim. Leaflet.Sim is a framework for location-based simulations with Leaflet maps that can visualise moving markers, which can change their style, and events over time on a map.
2emnlp2022-papers. Python
2finetune-evaluation-harness. Python
2turkish-lm-bias. Investigating Gender Bias in Turkish Language Models
2CmdLineSlideShow. Command line script for generating rich slide shows from a set of images with transition effects and audio. Using ImageMagick and FFMPEG.
2chat-ui. Open source codebase powering the HuggingChat app
1kibana-reallybettermap. Multiple locations for Kibana's bettermap panel
1Wikipedia2Lucene. Import a Wikipedia XML Dump from HDFS to Lucene index or Elasticsearch and retrieve similar Wikipedia articles based on Lucene's MoreLikeThis query.
1NeMo. NeMo: a toolkit for conversational AI
1data-sourcing. Python
1news-visualization. News visualization with Elastic Search and Kibana including NER, Sentiment Analysis and Geo Locations.
1