Hong Kong

Yufeng He

Intermediate
@he-yufeng

AI Researcher @ Moonshot AI (Kimi) | MS CS @ HKU | Champion, Shanghai Global AI Contest | 3× ACM-ICPC Silver | Former Intern @ Baidu, Maimai, Kuaishou

CoreCoder. Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. Think NanoGPT for coding agents. Formerly NanoCoder.

1.6k

FindJobs-Agent. LLM-powered toolkit for skill analysis, AI interviews, resume scoring, and job structuring. Automates professional skill taxonomy and interview processes with adaptive difficulty.

254

RepoWiki. Open-source DeepWiki alternative — generate comprehensive wiki documentation for any codebase from terminal or browser

221

ContractGuard. AI agent that reads the fine print so you don't have to. Upload any contract → get red flags, unfair terms, and plain-English explanations in seconds.

179

GitSense. AI-powered open source contribution finder and repo radar

102

AgentProbe. Drop-in pytest plugin for regression-testing AI agents — snapshot baselines, semantic comparison, mock LLMs

41

AnyCoder. AI coding agent in your terminal. Works with any LLM - DeepSeek, Qwen, GPT-5, Claude, Gemini, Kimi, Ollama. 100+ models via litellm. ~1300 lines of Python.

33

CodeABC. AI-powered code reader for non-programmers | 码上懂 — 不用学编程,也能读懂代码

11

PromptDiff. Semantic diff for LLM prompts — compare prompt versions like git diff

9

RuleForge. Auto-generate AI assistant rules (CLAUDE.md, .cursorrules, copilot-instructions) from codebase analysis

8

LiteBench. A pip-installable benchmark runner for LLMs and agents. Five minutes to your first eval.

8

CodeJoust. Pit AI coding agents against the same bug. Score them on tests, diff, cost, and time — pick the winning patch.

8

TokenTracker. Drop-in LLM cost tracker — change one import line, see where your money goes. Supports OpenAI, OpenRouter, Azure, Ollama.

7

FlightBox. Record, replay, and diff every LLM call your AI agent makes

7

gsm8k-prompt-engineering. Prompt-engineering study on LLM math reasoning (GSM8K) and code generation (HumanEval): zero/few-shot, self-consistency, self-verification, and experiments on prompt quality, complexity, demonstrations, and diversity.

5

agentcikit. CI-grade evidence and safety tools for AI agents, MCP servers, and upstream contributions

5

IslandEscape. 2D pixel-art survival trading game — four LLM-powered AI agents compete to escape the island

5

arxiv-paper-coder. Multi-agent system that turns a natural-language brief into a working codebase — planner, coder, and reviewer coordinated by an orchestrator with dependency tracking. Ships an arXiv daily-papers site demo.

4

adversarial-refinement-imputation. Companion code for the MiLeTS 2026 paper "When Does Adversarial Refinement Help?" — adapting R3GAN (NeurIPS 2024) to multivariate time series imputation. A clearly-scoped negative result and open problem.

4

BatchLLM. Batch processing for LLM APIs - CSV/JSONL in, processed out. Concurrent requests, retries, checkpointing, cost tracking.

4

IssueBenchKit. Turn a real GitHub issue into a repeatable coding-agent benchmark task

4

DRL-MultiFactorTrading. Deep Reinforcement Learning trading strategies: Double DQN with Transformer Attention + Multi-Factor Model (Fama-French inspired). Features adaptive risk management and volatility targeting.

3

he-yufeng. Profile README

3

TrajBias. TrajBias: Structural Biases in LLM-as-Judge Evaluation of Agent Trajectories

2

PatchContext. Python

2

ActionRepro. Python

2

he-yufeng.github.io. Yufeng He's personal blog

1

universeplayer.github.io. Yufeng He's personal blog - AI Engineer & ACM-ICPC Medalist

1

MCPReady. CI gate for MCP servers

1

MCPReplay. Python

1

d2l-zh. 《动手学深度学习》:面向中文读者、能运行、可讨论。中英文版被70多个国家的500多所大学用于教学。

1
31
Apply