Bruno Magalhaes

Machine Learning and High Performance Computing

linkedin github scholar resume email RSS
photo

Hi👋🏽! I am Bruno, an ML Systems researcher at Huawei Research Switzerland. Previously, I was an ML researcher at Microsoft Research Cambridge, and an HPC engineer, PhD, and postdoc at EPFL. My research focuses on improving edge-to-cloud ML system efficiency — from efficient on-device inference to thousand-GPU distributed training — for Transformer LLMs, MoE, diffusion models, and AI agents.

In this space, I post about past projects and topics related to my fields of interest:

2024 Distributed GPT model (part 4): sequence and context parallelism with Ulysses and Ring attention
2024 Distributed training of variable-length samples: curriculum learning, compilation, adaptive batch size and LR
2024 Distributed Mixture-of-Experts and Expert Parallelism
2023 Distributed GPT model (part 3): Megatron-LM tensor parallelism
2023 Distributed GPT model (part 2): pipeline parallelism
2023 Distributed GPT model: data parallelism, sharding and CPU offloading
2023 Building a GPT model in C++, and benchmarking LibTorch, PyTorch, TorchScript and torch.compile
2023 Building a GPT model in PyTorch from scratch
2020 Learning from sequences: Encoder-Decoder, Transformers and BERT
2019 Variational Autoencoders (VAEs) and Generative Adversarial Neural Networks (GANs)
2019 Variational Inference: ELBO, Mean-Field Approximation, CAVI and Gaussian Mixture Models
2019 Exponential Family of Distributions
2018 Bayesian Linear Regression, Maximum Likelihood and Maximum-A-Priori
2018 Statistics for ML Engineers
2018 Algebra for ML Engineers
2018 Deep Neural Networks, backpropagation, autodiff, dropout, CNNs and embeddings
2017 Unsupervised Learning basics and Principal Component Analysis
2017 Variable Timestep Simulation of the Electrical Activity of Neurons
2017 Closed-form Linear Regression and Matrix Factorization, and loss functions
2016 Numerical Resolution of the Electrical Activity of Detailed Neuron Models
2016 The Leaky Integrate-and-Fire Neuron Model and The Brunel Network
2015 Distributed Orthogonal Slicing for Load Balancing of Large Spatial Datasets
2015 Distributed Matrix Transpose Algorithms
2014 Distributed Sorting Algorithms

I also keep track of books and other related resources that are available online:

book Deep Learning - Foundations and Concepts, Christopher M. Bishop, Hugh Bishop (pdf)
book Pattern Classification and Machine Learning, Christopher M. Bishop (pdf)
book Model-Based Machine Learning, John Winn et al. (pdf)
book Mathematics for Machine Learning, Deisenroth, Aldo Faisal, Cheng S. Ong (pdf)
book Neuronal Dynamics, Wulfram Gerstner et al. (online)
book The Little Book of Deep Learning, Francois Fleuret (pdf)
book Understanding Machine Learning: From Theory to Algorithms, Shai Shalev-Shwartz and Shai Ben-David (pdf)
book Information Theory, Inference, and Learning Algorithms, David MacKay (pdf)
article In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
post The vLLM MoE Playbook: A Practical Guide to TP, DP, PP and Expert Parallelism, AMD
post How to scale your model, Jax ML
post The Ultra-Scale Playbook: Training LLMs on GPU Clusters, huggingface
notes The Matrix Cookbook
lecture Computing Gradients with Backpropagation - Automatic Differentiation (Princeton COS-324) (pdf)
lecture Statistics for data Science (EPFL MATH-413):

And finally, I keep a curated list of summaries for relevant publications. Enjoy, and feel free to reach out with bugs, improvements, or questions 🚀!