A thread-safe Swift wrapper for llama.cpp, designed with proper concurrency handling from the ground up. LlamaSwift provides a modern, async/await-based API for running large language models on macOS and iOS with full support for Apple Silicon GPU acceleration via Metal.
- ✅ Thread-safe: All llama.cpp operations run on a single serial queue
- ✅ Actor-based: Uses Swift actors for isolation
- ✅ Async/await: Modern Swift concurrency support
- ✅ Streaming: Token-by-token generation with async sequences
- ✅ Memory safe: Proper resource management
- ✅ GPU Acceleration: Metal backend support for Apple Silicon
- ✅ Multi-platform: Supports macOS 13+ and iOS 16+
- macOS 13.0+ or iOS 16.0+
- Xcode 15.0+
- Swift 5.9+
- llama.cpp (included as source)
Add LlamaSwift to your project using Swift Package Manager:
- In Xcode, go to File → Add Package Dependencies...
- Enter the repository URL:
https://github.com/yourusername/LlamaSwift.git - Select the version or branch you want to use
- Add
LlamaSwiftto your target's dependencies
Alternatively, add it to your Package.swift:
dependencies: [
.package(url: "https://github.com/yourusername/LlamaSwift.git", from: "1.0.0")
]import LlamaSwift
// Load a model
let model = try await LlamaModel.load(from: "/path/to/model.gguf")
// Generate text with streaming
let stream = try await model.generate(prompt: "Hello, how are you?")
for try await token in stream {
print(token, terminator: "")
}import LlamaSwift
// Load model with custom parameters
let model = try await LlamaModel.load(
from: "/path/to/model.gguf",
contextSize: 4096,
useGPU: true // Enable Metal GPU acceleration
)
// Generate with custom parameters
let stream = try await model.generate(
prompt: "Write a story about",
maxTokens: 100,
temperature: 0.7,
topP: 0.9
)
// Process tokens as they arrive
for try await token in stream {
// Handle each token
processToken(token)
}The wrapper ensures thread safety by:
- Using a single serial
DispatchQueuefor all llama.cpp operations - Wrapping the API in an
Actorfor Swift-level isolation - Never allowing concurrent access to llama.cpp structures
The wrapper bridges to llama.cpp through:
- C Interface (
llama_bridge.h): C functions that wrap llama.cpp - Objective-C++ Bridge (
llama_bridge.mm): Implements the C interface using llama.cpp - Swift API (
LlamaModel.swift): Swift-friendly async/await API
- CPU: Optimized CPU backend using Accelerate framework
- Metal: GPU acceleration on Apple Silicon devices
The main entry point for loading and using models.
load(from:contextSize:useGPU:)- Load a model from a file pathgenerate(prompt:maxTokens:temperature:topP:)- Generate text with streaming support
Contributions are welcome! Please feel free to submit a Pull Request. For major changes, please open an issue first to discuss what you would like to change.
[Add your license here - e.g., MIT, Apache 2.0, etc.]
- Built on top of llama.cpp by Georgi Gerganov