Building a DIY Yoga Pose Detection System

Most people who try to build a yoga pose detection game from scratch end up abandoning it around the pose validation stage. The pose detection part works fine in tutorials. The actual scoring logic is where things get complicated, and honestly, most open-source implementations skip over that properly. I built something like this a while back for a personal project. What I'm about to describe is the actual working approach, not the simplified version you see in blog posts.

Core Diy Yoga Pose Gameplay Architecture

The system breaks into three layers. First, you need skeleton extraction from video input. Second, you need angle and spatial reasoning to validate poses. Third, you need a scoring and feedback loop that actually feels fair to a human player. Most tutorials stop at layer one and call it a day. For skeleton extraction, I used MediaPipe Pose. It gives you 33 landmarks in real time on a decent CPU. That is enough for most basic yoga poses. You pull the coordinates for shoulders, hips, elbows, knees, and ankles. From there, you calculate joint angles using basic dot product math between limb vectors. The angle calculation part sounds straightforward. Take the elbow joint. You have the shoulder-to-elbow vector and the elbow-to-wrist vector. You compute the angle between them. If it is close enough to 180 degrees, the arm is straight. That works for simple cases. Yoga is not simple cases. Tree pose requires balance information that skeleton data alone cannot provide. Warrior II requires hip width and knee alignment that a single camera struggles with if the person is not facing directly forward.

Practical Scoring Logic

Here is where most implementations fail. They set a static tolerance range for each pose. Hold your arms at 90 degrees plus or minus ten degrees and you pass. This does not work in practice because people move. Even when holding a pose, there is natural sway. A rigid threshold makes the game feel unfair and causes constant false failures. I ended up implementing a rolling average score combined with a minimum hold time. The system samples pose angles at sixty frames per second and maintains a moving average over approximately two seconds. A pose only registers as complete when the average stays within tolerance for a sustained period. This is how I solved the edge case where someone enters a pose correctly but wobbles for the first second before settling. The rolling average ignores that initial instability without requiring the user to hold perfectly still from the moment they enter frame. Another issue I ran into was mirror reflections in the detection area. My testing room had a large window that reflected the subject back into the camera at certain times of day. The pose estimator occasionally detected two skeletons and assigned landmark data from both people mixed together. This produced impossible joint configurations that looked like the user had six limbs. The workaround was a confidence filter. If the model's confidence score for any landmark dropped below a threshold, I dropped that landmark entirely from the calculation rather than using bad data. A pose with insufficient valid landmarks simply did not register instead of registering incorrectly.

Get the Full Details

Wii Fit U - Yoga Dance Pose Gameplay - YouTube
Wii Fit U - Yoga Dance Pose Gameplay - YouTube

Implementation Notes

If you are building this yourself, start with a fixed camera position. Do not attempt to handle arbitrary camera movement. The math gets significantly harder and the accuracy drops regardless. Mount your device on a tripod or stack of books. Distance matters more than most people realize. At roughly eight to ten feet from the camera, MediaPipe Pose gives its most consistent full-body detection. Closer than that and the upper body landmarks remain accurate while the feet become unreliable. Farther than that and the whole skeleton loses precision. Lighting should be even and diffuse. Backlighting from a bright window behind the subject will cause landmark jitter because the contrast reversal confuses the depth estimation in cheaper webcams. A front-lit or side-lit setup works fine. Natural light from a window in front of you, not behind you, is ideal.

Known Limitations

This approach works for solo practice with basic to intermediate yoga poses. It will not handle partner poses, complex arm balances, or poses that require precise foot placement relative to the ground plane. The single-camera setup simply cannot determine lateral distance accurately enough for those. If you need that level of precision, you are looking at either a depth-sensing camera like an Intel RealSense or a second synchronized camera with stereo calibration, and that moves the project into a completely different complexity tier. The scoring is also inherently approximate. A human yoga instructor can detect a half-degree hip rotation that shifts weight distribution. Your algorithm cannot. This means the system will occasionally approve poses that a teacher would correct and reject poses that are technically acceptable. This is a hardware limitation, not a software bug. Accept it and design your feedback messages to be encouraging rather than clinical. For a practical starting point, check out the Pose estimation libraries on GitHub and build outward from there. The official Mediapipe Pose examples give you the skeleton pipeline in about an hour. Adding the yoga-specific scoring logic is where the real work begins.