This is your work, valued

Berlin, Germany

malteos

Expert
@malteos

Research engineer: Datasets, information retrieval, representation learning, LLMs, scientific & legal document processing

awesome-document-similarity. A curated list of resources on document similarity measures (papers, tutorials, code, ...)

256

pytorch-bert-document-classification. Enriching BERT with Knowledge Graph Embedding for Document Classification (PyTorch)

161

scincl. Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings (EMNLP 2022 paper)

79

llm-datasets. A collection of datasets for language model pretraining including scripts for downloading, preprocesssing, and sampling.

66

aspect-document-similarity. Implementation, trained models and result data for the paper "Aspect-based Document Similarity for Research Papers" #COLING2020

63

legal-document-similarity. Legal document similarity - Code, data, and models for the ICAIL 2021 paper "Evaluating Document Representations for Content-based Legal Literature Recommendations"

32

semantic-document-relations. Implementation, trained models and result data for the paper "Pairwise Multi-Class Document Classification for Semantic Relations between Wikipedia Articles"

31

clp-transfer. Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning

30

awesome-prompt-optimization. A curated collection of resources for prompt engineering, optimization, and automatic prompt generation across text, image, video, and multimodal AI systems.

18

awesome-anonymization-for-llms. A collection of resources for PII detection, anonymization, privacy-preserving techniques, and GDPR compliance in Large Language Model (LLM) or AI applications.

18

german-language-models. A collection of German GPT language models

11

aspect-document-embeddings. Code, dataset & models for the paper Specialized Document Embeddings for Aspect-based Similarity of Research Papers (#JCDL2022)

11

awesome-contrastive-learning-for-nlp. A collection of papers about contrastive learning for natural language processing.

7

wikipedia-article-recommendations. Survey data and Python code for the ICADL 2021 paper "A Qualitative Evaluation of User Preference for Link-based vs. Text-based Recommendations of Wikipedia Articles"

5

getting-started. Dockerfile

4

covid-vaccination-appointment. Python

3

Leaflet.Sim. Leaflet.Sim is a framework for location-based simulations with Leaflet maps that can visualise moving markers, which can change their style, and events over time on a map.

2

emnlp2022-papers. Python

2

finetune-evaluation-harness. Python

2

turkish-lm-bias. Investigating Gender Bias in Turkish Language Models

2

CmdLineSlideShow. Command line script for generating rich slide shows from a set of images with transition effects and audio. Using ImageMagick and FFMPEG.

2

chat-ui. Open source codebase powering the HuggingChat app

1

kibana-reallybettermap. Multiple locations for Kibana's bettermap panel

1

Wikipedia2Lucene. Import a Wikipedia XML Dump from HDFS to Lucene index or Elasticsearch and retrieve similar Wikipedia articles based on Lucene's MoreLikeThis query.

1

NeMo. NeMo: a toolkit for conversational AI

1

data-sourcing. Python

1

news-visualization. News visualization with Elastic Search and Kibana including NER, Sentiment Analysis and Geo Locations.

1