Poker RL: Building with GRPO
I wanted to see if I could take a small language model and teach it to play poker using reinforcement learning with verifiable rewards (RLVR). The idea is pretty simple: poker gives you a clear rew...
I wanted to see if I could take a small language model and teach it to play poker using reinforcement learning with verifiable rewards (RLVR). The idea is pretty simple: poker gives you a clear rew...
I read a decent amount of AI/ ML papers and wanted a way to consume on my way to work. I found that tools like NotebookLM that generate podcast-style audio tend to stay surface-level and I wanted s...
I wanted to see if I could track tennis players from broadcast footage without labeling a single frame. Most sports tracking projects start with collecting labeled data and training a YOLO model, b...
This is the final part of the series where we actually render an image. Part 1 covered projecting 3D points to 2D, Part 2 covered the Gaussian math and covariance projection. Now we take all of tha...
Prompt engineering is a tedious and frankly annoying task: write a prompt, check the output, tweak the wording, and repeat without ever being confident it’ll generalize to thousands of situations. ...
In Part 1 we covered how to take 3D points from a COLMAP scan and project them onto a 2D image plane. Part 2 gets into the actual Gaussian part of Gaussian Splatting: how each splat is represented,...
I wanted to understand 3D Gaussian Splatting from the ground up, but I kept running into the same problem: almost every implementation is written in CUDA, and most explanations assume you already k...
Analytics have completely changed basketball but the tooling is expensive. Companies like Second Spectrum give NBA teams detailed breakdowns of shot quality, defensive coverages, and more, but most...
We collected machine learning final exams from MIT, Harvard, and Cornell and threw GPT-3, OPT, Codex, and ChatGPT at them. These weren’t just single answer problem sets. They’re multi-part, multi-t...