A multi-backend desktop application for human pose estimation across images, webcam streams, and video files
A local-first computer vision application built with React, Tauri, Rust, and Python, supporting both MediaPipe Pose Landmarker and YOLO Pose through a unified pose-estimation engine.
Overview β’ Features β’ Architecture β’ Backends β’ Installation β’ Development β’ Testing
Human Pose Estimation is a local Windows desktop application for estimating human body poses from:
- Still images
- Live webcam input
- Local video files
The application combines a modern React interface with a Tauri/Rust desktop layer and a persistent Python computer-vision engine.
Instead of coupling the interface directly to one pose-estimation library, the project uses a backend-independent pose contract that allows different inference engines to expose results through one normalized representation.
Currently supported pose backends include:
- MediaPipe Pose Landmarker
- Ultralytics YOLO Pose
All inference runs locally on the user's machine. The application does not rely on a web server or cloud inference service.
- πΌοΈ Human pose estimation from still images
- π· Live webcam pose estimation
- π¬ Pose analysis for local video files
- π§ Multiple pose-estimation backends
- π§ Multi-person pose support
- 𦴠Skeleton and keypoint visualization
- π¦ Optional person bounding boxes
- π Backend-independent pose representation
- β‘ Persistent Python inference process
- π Typed Rust β Python communication
- π΄ Offline-friendly runtime behavior
- π Custom local model selection
- π§© Explicit model-asset handling
- πͺ Native Windows desktop packaging
- π§ͺ Python, frontend, and Rust test suites
Load a supported image and estimate one or more human poses.
The interface can display:
- keypoints
- skeleton connections
- bounding boxes
- detected person count
- image dimensions
- inference backend
- processing time
The application supports live webcam estimation using bounded frame processing.
Only one inference request is kept active at a time, preventing stale frames from building up when inference is slower than the webcam preview.
Local videos can be played using native video controls while pose analysis samples fresh frames during playback.
Old results are discarded after seeking or when they fall too far behind the current playback position.
The system separates the desktop interface from the computer-vision engine.
βββββββββββββββββββββββββββββββββ
β Input Source β
β β
β Image Webcam Video β
βββββββββββββββββ¬ββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββ
β React UI β
β β
β Preview β’ Controls β’ Overlay β
βββββββββββββββββ¬ββββββββββββββββ
β
Tauri Commands
β
βΌ
βββββββββββββββββββββββββββββββββ
β Rust / Tauri β
β β
β Sidecar Process Manager β
β Request IDs β’ Timeouts β
βββββββββββββββββ¬ββββββββββββββββ
β
NDJSON IPC
β
βΌ
βββββββββββββββββββββββββββββββββ
β Python Sidecar β
β β
β PoseEngine β
βββββββββββββ¬ββββββββββ¬ββββββββββ
β β
βΌ βΌ
MediaPipe YOLO Pose
β β
ββββββ¬βββββ
βΌ
Unified PoseResult
β
βΌ
SVG Pose Overlay
The project uses a persistent sidecar architecture rather than starting a Python process for every inference request.
This design provides several advantages:
- model instances can be reused
- repeated Python startup overhead is avoided
- communication remains explicit and testable
- frontend code stays independent of backend-specific inference objects
- long-running webcam and video analysis remains memory-bounded
Communication between Rust and Python uses line-delimited JSON over standard process streams.
| Backend | Status | Output |
|---|---|---|
| MediaPipe Pose Landmarker | β Supported | Up to 33 canonical landmarks |
| YOLO Pose | β Supported | COCO 17-keypoint pose observations, including multi-person output |
| MMPose | Current Python 3.13 / Windows dependency stack is not reproducible |
Both supported backends are normalized into the same application-level pose contract.
This means the React interface and Rust bridge do not need backend-specific rendering logic.
One of the key design decisions in this project is the use of a backend-independent pose model.
Instead of exposing raw MediaPipe or Ultralytics objects to the application, each backend is adapted into a common PoseResult structure.
Conceptually:
MediaPipe Result ββ
β
ββββΊ PoseResult ββββΊ Application
β
YOLO Pose Result ββ
This makes backend selection transparent to the rest of the application.
| Layer | Technology |
|---|---|
| Interface | React 19, Vite 7, JavaScript |
| Desktop Shell | Tauri 2 |
| Native Bridge | Rust |
| Vision Engine | Python 3.13.5 |
| Image Processing | OpenCV, NumPy |
| Pose Backend | MediaPipe Tasks |
| Pose Backend | Ultralytics YOLO Pose |
| IPC | Persistent NDJSON over process streams |
| Packaging | PyInstaller, Tauri, NSIS |
| Target Platform | Windows x64 |
Human-Pose-Estimation/
β
βββ src/
β βββ React user interface, overlays, and media schedulers
β
βββ src-tauri/
β βββ Rust bridge, Tauri commands, sidecar management, and packaging
β
βββ python-engine/
β βββ Pose backends, PoseEngine, protocol, serialization, and tests
β
βββ scripts/
β βββ Development and Windows release automation
β
βββ index.html
βββ package.json
βββ vite.config.js
βββ yarn.lock
βββ .gitignore
βββ README.md
Model weights are intentionally kept external to the repository.
The application does not silently download pose models.
Default development location:
python-engine/models/mediapipe/pose_landmarker.task
Default development location:
python-engine/models/yolo/yolo11n-pose.pt
Users can also select compatible local model files through the application interface.
The packaged application targets:
Windows 11 x64
The release build includes the Python runtime and required Python dependencies.
End users therefore do not need to separately install:
- Python
- pip
- Node.js
- Rust
- Tauri
The current installer is generated using NSIS.
The installer is currently unsigned, so Windows SmartScreen may display a warning.
For source development:
- Windows 11 x64
- Node.js
- Yarn 1.x
- Rust toolchain
- Tauri Windows prerequisites
- Python 3.13.5
From the repository root:
python -m venv python-engine\.venvInstall Python dependencies:
.\python-engine\.venv\Scripts\python.exe -m pip install --upgrade pip
.\python-engine\.venv\Scripts\python.exe -m pip install -e ".\python-engine[test]"Install frontend dependencies:
yarn installRun the desktop application in development mode:
yarn tauri devInstall the packaging dependencies:
.\python-engine\.venv\Scripts\python.exe -m pip install -e ".\python-engine[package]"Run the Windows release build:
.\scripts\build-windows-release.ps1The build process:
- verifies the expected Python runtime
- packages the Python sidecar
- smoke-tests the packaged sidecar
- embeds the runtime into the Tauri application
- builds the frontend
- builds the Rust desktop shell
- creates an NSIS Windows installer
Generated installers are placed under:
src-tauri/target/release/bundle/nsis/
The repository contains separate test suites for Python, frontend, and Rust components.
yarn test
yarn buildcd python-engine
.\.venv\Scripts\python.exe -m pytest
.\.venv\Scripts\python.exe -m pytest -m integration -racd src-tauri
cargo fmt --check
cargo check --locked
cargo test --locked
cargo clippy --locked -- -D warningsThe latest verified project state includes:
243 Python tests passed
43 frontend tests passed
23 active Rust tests passed
The packaged Python sidecar smoke test and packaged-runtime protocol checks also pass.
Some of the main engineering decisions behind the project include:
Python remains alive between requests, allowing initialized pose estimators to be reused.
MediaPipe and YOLO Pose are hidden behind one application-level estimator contract.
Structured error codes are preserved across:
Python
β
NDJSON
β
Rust
β
React
while the UI presents safe, readable guidance to users.
Webcam and video analysis keep at most one inference request active.
This prevents unbounded frame queues and stale inference results.
Model downloads never happen implicitly.
This keeps the application deterministic and makes offline usage possible.
- Windows x64 is currently the only packaged target
- The Windows installer is unsigned
- Production packaging currently uses CPU-only Torch
- Model files must be provided separately
- MMPose is currently unavailable in the supported Python/Windows environment
- Temporal tracking and pose smoothing are not implemented
- Action recognition is not implemented
- Processed video export is not currently supported
Potential future improvements include:
- Project-specific application icon
- Signed Windows release
- Public downloadable installer
- GPU-enabled production package
- Cross-platform desktop builds
- Temporal pose smoothing
- Multi-frame person tracking
- Action recognition
- Processed video export
- Additional pose-estimation backends
- Revisit MMPose when dependency support improves
Real application screenshots should be added to:
docs/images/image-mode.png
docs/images/webcam-mode.png
docs/images/video-mode.png
Recommended README layout:
### Image Mode

### Webcam Mode

### Video Mode
Using real release screenshots is preferable to using mockups because it shows the actual application state.
This project explores both the computer-vision and software-engineering aspects of deploying pose-estimation models in a desktop application.
It demonstrates:
- human pose estimation
- multi-backend computer vision
- model abstraction
- local process communication
- desktop application architecture
- live-media processing
- Rust/Python integration
- React visualization
- Windows application packaging
- automated testing across multiple technology stacks
Contributions, suggestions, and issue reports are welcome.
- Fork the repository
- Create a feature branch
git checkout -b feature/your-feature- Commit your changes
git commit -m "Add new feature"- Push the branch
git push origin feature/your-feature- Open a Pull Request