This is your work, valued

Mariusz Kurman

Expert
@mkurman

zorai. Zorai is a persistent, multi-agent, auditable, learning execution platform where the daemon owns work, memory, approvals, tools, and long-running goals.

318

cmux-windows. Windows terminal with useful features. Fully vibe-coded using CC and Codex.

258

synthetic-questions-generation. Python

88

grpo-llm-evaluator. Fine-tunes a student LLM using teacher feedback for improved reasoning and answer quality. Implements GRPO with teacher-provided evaluations.

54

neuroblast-v3. NeuroBLAST v3 architecture code

37

synthlabs. Create synthetic datasets from scratch using AI-powered generation. Define topics, customize prompts, and generate high-quality reasoning traces in the SYNTH format.

32

ReasonFlow. ReasonFlow is a novel framework designed to implement o1-like reasoning capabilities in large language models.

19

mcts-pytorch. A flexible Monte Carlo Tree Search framework with PyTorch for decision-making in language models.

12

convgpt. Python

12

synthetic-conversation-generator. A Python script to programmatically generate synthetic conversations for training LLMs, chatbots, and dialogue systems.

8

blurred-thoughts-SFT. Blurred-Thoughts Supervised-Finetuning (BT-SFT) is a new approach to fine-tuning language models, focusing on enhancing response diversity and creativity.

7

jepa-llm. Fine-tuning causal language models with an additional JEPA-style representation regularisation loss as well as a plain Hugging Face trainer

6

Large-Language-Model-Notebooks-Course. Practical course about Large Language Models.

2

fixed-size-kv-cache. This project implements a fixed-size key-value cache for use in transformer models, specifically designed to work with the LLaMA model. The cache dynamically truncates the key and value states based on attention weights, ensuring efficient memory usage and performance.

2

self_reward_head_pytorch. This repository contains the implementation of a self-reward head designed for language models. The self-reward head enables the model to autonomously score its generated outputs, promoting self-assessment and iterative improvement.

2

nvamp-loss. Normalized Variance-Aware Max-Penalized Loss

1

linearmoe_pytorch. This repo contains my custom implementation of a mixture of experts as an extension of the linear layer.

1

unsloth-studio-docker. Shell

1