I build and test local AI systems on Apple silicon and NVIDIA hardware. My work covers MLX, CUDA, mixed-hardware inference, local agents and runtime security.
The interesting part is not prompting a model.
The interesting part is building the system around it securely, observably and in production.
Build logs on X · Models on Hugging Face
Vontra · local inference enablement
Helping make local inference more accessible through MLX model creation, testing, validation, and distribution.
TensorFold · Apple silicon / MLX local inference
Local-first runtime work for running sparse MoE language models on Apple silicon with MLX under tight memory budgets, with bounded resident memory, explicit paging, and runtime telemetry.




