Rare find

FlexLLMGen. Run large AI models on a single graphics card without slowing down.

github.com/FMInference/FlexLLMGen

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

October 2024
  • Rename for compliance (#141)
  • Update profile_matmul.py
  • Update profile_matmul.py
September 2024
  • Update README.md (#139)
  • Update profile_bandwidth.py
April 2024
  • upload a draft script for fitting the cost model
July 2023
  • upload poster
  • Update README.md
  • Add cost model (#121)
June 2023
  • Add instructions for running the Petals benchmarks
April 2023
  • Update README with more instructions (#110)
March 2023
  • Update README.md
  • Update paper.md (#102)
  • Data wrangle benchmark (#95)
  • Update README.md
  • Update Petals setup details
  • Update links in README
  • Update Petals results in README
  • Update links
  • Update README.md