distributed-llama. Tensor parallelism is all you need. Run LLMs on weak devices or make powerful devices even more powerful by distributing the workload and dividing the RAM usage.

github.com/adamcohenhillel/distributed-llama

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.