Skip to content

Latest commit

Β 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🧍 Human Pose Estimation

A multi-backend desktop application for human pose estimation across images, webcam streams, and video files

A local-first computer vision application built with React, Tauri, Rust, and Python, supporting both MediaPipe Pose Landmarker and YOLO Pose through a unified pose-estimation engine.


Python React Tauri Rust MediaPipe YOLO


Overview β€’ Features β€’ Architecture β€’ Backends β€’ Installation β€’ Development β€’ Testing


πŸ“Œ Overview

Human Pose Estimation is a local Windows desktop application for estimating human body poses from:

  • Still images
  • Live webcam input
  • Local video files

The application combines a modern React interface with a Tauri/Rust desktop layer and a persistent Python computer-vision engine.

Instead of coupling the interface directly to one pose-estimation library, the project uses a backend-independent pose contract that allows different inference engines to expose results through one normalized representation.

Currently supported pose backends include:

  • MediaPipe Pose Landmarker
  • Ultralytics YOLO Pose

All inference runs locally on the user's machine. The application does not rely on a web server or cloud inference service.


✨ Features

  • πŸ–ΌοΈ Human pose estimation from still images
  • πŸ“· Live webcam pose estimation
  • 🎬 Pose analysis for local video files
  • 🧠 Multiple pose-estimation backends
  • 🧍 Multi-person pose support
  • 🦴 Skeleton and keypoint visualization
  • πŸ“¦ Optional person bounding boxes
  • πŸ”„ Backend-independent pose representation
  • ⚑ Persistent Python inference process
  • πŸ” Typed Rust ↔ Python communication
  • πŸ“΄ Offline-friendly runtime behavior
  • πŸ“‚ Custom local model selection
  • 🧩 Explicit model-asset handling
  • πŸͺŸ Native Windows desktop packaging
  • πŸ§ͺ Python, frontend, and Rust test suites

🎬 Application Modes

Image Mode

Load a supported image and estimate one or more human poses.

The interface can display:

  • keypoints
  • skeleton connections
  • bounding boxes
  • detected person count
  • image dimensions
  • inference backend
  • processing time

Webcam Mode

The application supports live webcam estimation using bounded frame processing.

Only one inference request is kept active at a time, preventing stale frames from building up when inference is slower than the webcam preview.

Video Mode

Local videos can be played using native video controls while pose analysis samples fresh frames during playback.

Old results are discarded after seeking or when they fall too far behind the current playback position.


🧠 How It Works

The system separates the desktop interface from the computer-vision engine.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚          Input Source         β”‚
β”‚                               β”‚
β”‚   Image   Webcam   Video      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                β”‚
                β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚          React UI             β”‚
β”‚                               β”‚
β”‚ Preview β€’ Controls β€’ Overlay  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                β”‚
         Tauri Commands
                β”‚
                β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚        Rust / Tauri           β”‚
β”‚                               β”‚
β”‚   Sidecar Process Manager     β”‚
β”‚ Request IDs β€’ Timeouts        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                β”‚
          NDJSON IPC
                β”‚
                β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚      Python Sidecar           β”‚
β”‚                               β”‚
β”‚         PoseEngine            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
            β”‚         β”‚
            β–Ό         β–Ό
      MediaPipe      YOLO Pose
            β”‚         β”‚
            β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
                 β–Ό
        Unified PoseResult
                 β”‚
                 β–Ό
          SVG Pose Overlay

πŸ—οΈ Architecture

The project uses a persistent sidecar architecture rather than starting a Python process for every inference request.

This design provides several advantages:

  • model instances can be reused
  • repeated Python startup overhead is avoided
  • communication remains explicit and testable
  • frontend code stays independent of backend-specific inference objects
  • long-running webcam and video analysis remains memory-bounded

Communication between Rust and Python uses line-delimited JSON over standard process streams.


πŸ”Œ Supported Backends

Backend Status Output
MediaPipe Pose Landmarker βœ… Supported Up to 33 canonical landmarks
YOLO Pose βœ… Supported COCO 17-keypoint pose observations, including multi-person output
MMPose ⚠️ Unavailable Current Python 3.13 / Windows dependency stack is not reproducible

Both supported backends are normalized into the same application-level pose contract.

This means the React interface and Rust bridge do not need backend-specific rendering logic.


🧩 Unified Pose Representation

One of the key design decisions in this project is the use of a backend-independent pose model.

Instead of exposing raw MediaPipe or Ultralytics objects to the application, each backend is adapted into a common PoseResult structure.

Conceptually:

MediaPipe Result ─┐
                  β”‚
                  β”œβ”€β”€β–Ί PoseResult ───► Application
                  β”‚
YOLO Pose Result β”€β”˜

This makes backend selection transparent to the rest of the application.


πŸ› οΈ Tech Stack

Layer Technology
Interface React 19, Vite 7, JavaScript
Desktop Shell Tauri 2
Native Bridge Rust
Vision Engine Python 3.13.5
Image Processing OpenCV, NumPy
Pose Backend MediaPipe Tasks
Pose Backend Ultralytics YOLO Pose
IPC Persistent NDJSON over process streams
Packaging PyInstaller, Tauri, NSIS
Target Platform Windows x64

πŸ“ Project Structure

Human-Pose-Estimation/
β”‚
β”œβ”€β”€ src/
β”‚   └── React user interface, overlays, and media schedulers
β”‚
β”œβ”€β”€ src-tauri/
β”‚   └── Rust bridge, Tauri commands, sidecar management, and packaging
β”‚
β”œβ”€β”€ python-engine/
β”‚   └── Pose backends, PoseEngine, protocol, serialization, and tests
β”‚
β”œβ”€β”€ scripts/
β”‚   └── Development and Windows release automation
β”‚
β”œβ”€β”€ index.html
β”œβ”€β”€ package.json
β”œβ”€β”€ vite.config.js
β”œβ”€β”€ yarn.lock
β”œβ”€β”€ .gitignore
└── README.md

πŸ“¦ Model Setup

Model weights are intentionally kept external to the repository.

The application does not silently download pose models.

MediaPipe

Default development location:

python-engine/models/mediapipe/pose_landmarker.task

YOLO Pose

Default development location:

python-engine/models/yolo/yolo11n-pose.pt

Users can also select compatible local model files through the application interface.


πŸš€ Installation

Windows Application

The packaged application targets:

Windows 11 x64

The release build includes the Python runtime and required Python dependencies.

End users therefore do not need to separately install:

  • Python
  • pip
  • Node.js
  • Rust
  • Tauri

The current installer is generated using NSIS.

The installer is currently unsigned, so Windows SmartScreen may display a warning.


πŸ’» Development

Requirements

For source development:

  • Windows 11 x64
  • Node.js
  • Yarn 1.x
  • Rust toolchain
  • Tauri Windows prerequisites
  • Python 3.13.5

Python Environment

From the repository root:

python -m venv python-engine\.venv

Install Python dependencies:

.\python-engine\.venv\Scripts\python.exe -m pip install --upgrade pip
.\python-engine\.venv\Scripts\python.exe -m pip install -e ".\python-engine[test]"

Install frontend dependencies:

yarn install

Run the desktop application in development mode:

yarn tauri dev

🏭 Production Build

Install the packaging dependencies:

.\python-engine\.venv\Scripts\python.exe -m pip install -e ".\python-engine[package]"

Run the Windows release build:

.\scripts\build-windows-release.ps1

The build process:

  1. verifies the expected Python runtime
  2. packages the Python sidecar
  3. smoke-tests the packaged sidecar
  4. embeds the runtime into the Tauri application
  5. builds the frontend
  6. builds the Rust desktop shell
  7. creates an NSIS Windows installer

Generated installers are placed under:

src-tauri/target/release/bundle/nsis/

πŸ§ͺ Testing

The repository contains separate test suites for Python, frontend, and Rust components.

Frontend

yarn test
yarn build

Python

cd python-engine

.\.venv\Scripts\python.exe -m pytest
.\.venv\Scripts\python.exe -m pytest -m integration -ra

Rust / Tauri

cd src-tauri

cargo fmt --check
cargo check --locked
cargo test --locked
cargo clippy --locked -- -D warnings

The latest verified project state includes:

243 Python tests passed
43 frontend tests passed
23 active Rust tests passed

The packaged Python sidecar smoke test and packaged-runtime protocol checks also pass.


βš™οΈ Engineering Highlights

Some of the main engineering decisions behind the project include:

Persistent Python Sidecar

Python remains alive between requests, allowing initialized pose estimators to be reused.

Backend Abstraction

MediaPipe and YOLO Pose are hidden behind one application-level estimator contract.

Typed Error Propagation

Structured error codes are preserved across:

Python
  ↓
NDJSON
  ↓
Rust
  ↓
React

while the UI presents safe, readable guidance to users.

Bounded Live Processing

Webcam and video analysis keep at most one inference request active.

This prevents unbounded frame queues and stale inference results.

Explicit Model Assets

Model downloads never happen implicitly.

This keeps the application deterministic and makes offline usage possible.


⚠️ Current Limitations

  • Windows x64 is currently the only packaged target
  • The Windows installer is unsigned
  • Production packaging currently uses CPU-only Torch
  • Model files must be provided separately
  • MMPose is currently unavailable in the supported Python/Windows environment
  • Temporal tracking and pose smoothing are not implemented
  • Action recognition is not implemented
  • Processed video export is not currently supported

πŸ—ΊοΈ Roadmap

Potential future improvements include:

  • Project-specific application icon
  • Signed Windows release
  • Public downloadable installer
  • GPU-enabled production package
  • Cross-platform desktop builds
  • Temporal pose smoothing
  • Multi-frame person tracking
  • Action recognition
  • Processed video export
  • Additional pose-estimation backends
  • Revisit MMPose when dependency support improves

πŸ“Έ Screenshots

Real application screenshots should be added to:

docs/images/image-mode.png
docs/images/webcam-mode.png
docs/images/video-mode.png

Recommended README layout:

### Image Mode

![Image Mode](docs/images/image-mode.png)

### Webcam Mode

![Webcam Mode](docs/images/webcam-mode.png)

### Video Mode

![Video Mode](docs/images/video-mode.png)

Using real release screenshots is preferable to using mockups because it shows the actual application state.


🎯 Project Purpose

This project explores both the computer-vision and software-engineering aspects of deploying pose-estimation models in a desktop application.

It demonstrates:

  • human pose estimation
  • multi-backend computer vision
  • model abstraction
  • local process communication
  • desktop application architecture
  • live-media processing
  • Rust/Python integration
  • React visualization
  • Windows application packaging
  • automated testing across multiple technology stacks

🀝 Contributing

Contributions, suggestions, and issue reports are welcome.

  1. Fork the repository
  2. Create a feature branch
git checkout -b feature/your-feature
  1. Commit your changes
git commit -m "Add new feature"
  1. Push the branch
git push origin feature/your-feature
  1. Open a Pull Request

πŸ‘€ Author

Moien Sohani Darban

GitHub


⭐ If you find this project useful, consider giving it a star.

Built with React, Tauri, Rust, Python, MediaPipe, and YOLO Pose

About

Desktop human pose estimation app for images, webcam, and video using MediaPipe, YOLO Pose, React, Tauri, Rust, and Python.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages