Bruno Magalhaes
Machine Learning and High Performance Computing
Hi👋🏽! I am Bruno, an ML Systems researcher at Huawei Research Switzerland. Previously, I was an ML researcher at Microsoft Research Cambridge, and an HPC engineer, PhD, and postdoc at EPFL.
My research focuses on improving edge-to-cloud ML system efficiency — from efficient on-device inference to thousand-GPU distributed training — for Transformer LLMs, MoE, diffusion models, and AI agents.
In this space, I post about past projects and topics related to my fields of interest:
|
2024
|
Distributed GPT model (part 4): sequence and context parallelism with Ulysses and Ring attention
|
|
2024
|
Distributed training of variable-length samples: curriculum learning, compilation, adaptive batch size and LR
|
|
2024
|
Distributed Mixture-of-Experts and Expert Parallelism
|
|
2023
|
Distributed GPT model (part 3): Megatron-LM tensor parallelism
|
|
2023
|
Distributed GPT model (part 2): pipeline parallelism
|
|
2023
|
Distributed GPT model: data parallelism, sharding and CPU offloading
|
|
2023
|
Building a GPT model in C++, and benchmarking LibTorch, PyTorch, TorchScript and torch.compile
|
|
2023
|
Building a GPT model in PyTorch from scratch
|
|
2020
|
Learning from sequences: Encoder-Decoder, Transformers and BERT
|
|
2019
|
Variational Autoencoders (VAEs) and Generative Adversarial Neural Networks (GANs)
|
|
2019
|
Variational Inference: ELBO, Mean-Field Approximation, CAVI and Gaussian Mixture Models
|
|
2019
|
Exponential Family of Distributions
|
|
2018
|
Bayesian Linear Regression, Maximum Likelihood and Maximum-A-Priori
|
|
2018
|
Statistics for ML Engineers
|
|
2018
|
Algebra for ML Engineers
|
|
2018
|
Deep Neural Networks, backpropagation, autodiff, dropout, CNNs and embeddings
|
|
2017
|
Unsupervised Learning basics and Principal Component Analysis
|
|
2017
|
Variable Timestep Simulation of the Electrical Activity of Neurons
|
|
2017
|
Closed-form Linear Regression and Matrix Factorization, and loss functions
|
|
2016
|
Numerical Resolution of the Electrical Activity of Detailed Neuron Models
|
|
2016
|
The Leaky Integrate-and-Fire Neuron Model and The Brunel Network
|
|
2015
|
Distributed Orthogonal Slicing for Load Balancing of Large Spatial Datasets
|
|
2015
|
Distributed Matrix Transpose Algorithms
|
|
2014
|
Distributed Sorting Algorithms
|
I also keep track of books and other related resources that are available online:
| book |
Deep Learning - Foundations and Concepts, Christopher M. Bishop, Hugh Bishop (pdf)
|
| book |
Pattern Classification and Machine Learning, Christopher M. Bishop (pdf)
|
| book |
Model-Based Machine Learning, John Winn et al. (pdf)
|
| book |
Mathematics for Machine Learning, Deisenroth, Aldo Faisal, Cheng S. Ong (pdf)
|
| book |
Neuronal Dynamics, Wulfram Gerstner et al. (online)
|
| book |
The Little Book of Deep Learning, Francois Fleuret (pdf)
|
| book |
Understanding Machine Learning: From Theory to Algorithms, Shai Shalev-Shwartz and Shai Ben-David (pdf)
|
| book |
Information Theory, Inference, and Learning Algorithms, David MacKay (pdf)
|
| article |
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
|
| post |
The vLLM MoE Playbook: A Practical Guide to TP, DP, PP and Expert Parallelism, AMD
|
| post |
How to scale your model, Jax ML
|
| post |
The Ultra-Scale Playbook: Training LLMs on GPU Clusters, huggingface
|
| notes |
The Matrix Cookbook
|
| lecture |
Computing Gradients with Backpropagation - Automatic Differentiation (Princeton COS-324) (pdf)
|
| lecture |
Statistics for data Science (EPFL MATH-413):
|
And finally, I keep a curated list of summaries for relevant publications. Enjoy, and feel free to reach out with bugs, improvements, or questions 🚀!