Bruno Magalhaes

Machine Learning and High-Performance Computing

linkedin github scholar resume email RSS
photo

Hi👋🏽! I am Bruno, an ML Systems Researcher at Huawei Research Switzerland. Previously, I was an ML Researcher at Microsoft Research Cambridge, as well as an HPC Engineer, PhD, and Postdoc at EPFL. I specialize in edge-to-cloud ML system efficiency — from resource-constrained on-device inference to GPU distributed pre-training — for Large Language Models (LLMs), Mixture-of-Experts (MoEs), diffusion models, and AI agents.

In this space, I write about my projects and topics related to my fields of interest:

2024 Distributed GPT model (part 4): sequence and context parallelism with Ulysses and Ring attention
2024 Distributed training of variable-length samples: curriculum learning, compilation, adaptive batch size and LR
2024 Distributed Mixture-of-Experts and Expert Parallelism
2023 Distributed GPT model (part 3): Megatron-LM tensor parallelism
2023 Distributed GPT model (part 2): pipeline parallelism
2023 Distributed GPT model: data parallelism, sharding and CPU offloading
2023 Building a GPT model in C++, and benchmarking LibTorch, PyTorch, TorchScript and torch.compile
2023 Building a GPT model in PyTorch from scratch
2020 Learning from sequences: Encoder-Decoder, Transformers and BERT
2019 Variational Autoencoders (VAEs) and Generative Adversarial Neural Networks (GANs)
2019 Variational Inference: ELBO, Mean-Field Approximation, CAVI and Gaussian Mixture Models
2019 Exponential Family of Distributions
2018 Bayesian Linear Regression, Maximum Likelihood and Maximum-A-Priori
2018 Statistics for ML Engineers
2018 Algebra for ML Engineers
2018 Deep Neural Networks, backpropagation, autodiff, dropout, CNNs and embeddings
2017 Unsupervised Learning basics and Principal Component Analysis
2017 Variable Timestep Simulation of the Electrical Activity of Neurons
2017 Closed-form Linear Regression and Matrix Factorization, and loss functions
2016 Numerical Resolution of the Electrical Activity of Detailed Neuron Models
2016 The Leaky Integrate-and-Fire Neuron Model and The Brunel Network
2015 Distributed Orthogonal Slicing for Load Balancing of Large Spatial Datasets
2015 Distributed Matrix Transpose Algorithms
2014 Distributed Sorting Algorithms

I also keep a curated list of related books, articles, and reference materials available online:

book Deep Learning - Foundations and Concepts, Christopher M. Bishop, Hugh Bishop (pdf)
book Pattern Recognition and Machine Learning, Christopher M. Bishop (pdf)
book Model-Based Machine Learning, John Winn et al. (pdf)
book Mathematics for Machine Learning, Marc Peter Deisenroth, A. Aldo Faisal, Cheng Soon Ong (pdf)
book Neuronal Dynamics, Wulfram Gerstner et al. (online)
book The Little Book of Deep Learning, François Fleuret (pdf)
book Understanding Machine Learning: From Theory to Algorithms, Shai Shalev-Shwartz and Shai Ben-David (pdf)
book Information Theory, Inference, and Learning Algorithms, David MacKay (pdf)
article In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
post The vLLM MoE Playbook: A Practical Guide to TP, DP, PP and Expert Parallelism (AMD)
post How to scale your model (JAX ML)
post The Ultra-Scale Playbook: Training LLMs on GPU Clusters (Hugging Face)
notes The Matrix Cookbook
lecture Computing Gradients with Backpropagation - Automatic Differentiation (Princeton course COS-324) (pdf)
course Statistics for Data Science (EPFL course MATH-413)

And finally, I maintain a list of relevant publications. Enjoy, and feel free to reach out with feedback, fixes, or questions 🚀!