This is your work, valued
@ Cambridge CS PhD Candidate @ HKUST CSE RA
IJCAI2023-OptimalShardedDataParallel. [IJCAI2023] An automated parallel training system that combines the advantages from both data and model parallelism. If you have any interests, please visit/star/fork https://github.com/Youhe-Jiang/OptimalShardedDataParallel
52Hetu. A high-performance distributed deep learning system targeting large-scale and automated distributed training.
2AutoSharded_Transformer_Based_on_PyTorch. Python
2alpa. Training and serving large-scale neural networks with auto parallelization.
1Dynasor. Simple extension on vLLM to help you speed up reasoning model without training.
1apex. A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
1baidu-allreduce. Cuda
1annotation.
1DCT. A Conditional Independence Test in the Presence of Discretization
1bagua. Bagua is a deep learning training acceleration framework for PyTorch. It provides a one-stop training acceleration solution, including faster distributed training compared to PyTorch DDP, faster dataloader, kernel fusion, and more.
1ChooseTreeStructure. An algorithm to display tree structure
1AutoShard_Tool. a tool for recursive autoshard
1AllReduce-Over-MPI. C++
1BenchmarkDDP. Python
1AutoTopology-CostModelForOperators. Modeling and simulating experiments on communication and computing costs.
1ThunderServe. ThunderServe: Accelerating LLM Serving in Cloud Environments
1