
Hi, I’m Brayan a ML engineer based in Nairobi. I recently finished my B.Sc. in Information Technology at KCA University and have been spending most of my time building things at the intersection of reinforcement learning, language models, and low-level systems.
I built a full RLHF pipeline from scratch in PyTorch SFT, PPO, GRPO, and DPO and ran a cross-algorithm comparison study whose results are documented here. I also trained a minGPT in JAX/Flax on TinyStories when I was first learning the stack.
I write about what I’m building as I build it.
Apart from that I enjoy being outside and hanging out with friends, running and watching movies.
May 2026 Collaborating on VAML PPO + distributional value network for LLM RL
Posts
-
LLM From Scratch: Architecture and Training Infrastructure
-
From Pixels to Labels: Building a Satellite Image Segmentation Model
-
Value network architecture for llm reinforcement learning
-
Hardware Economics Blog
-
RLHF from Scratch: A Complete Alignment Study
-
Implementing Direct Preference Optimization (DPO)
-
Building a GPT from Scratch in JAX/Flax:
subscribe via RSS