Building the app with LLaMA 3.2 1B integration...
download_llama32_model.py- Downloads LLaMA 3.2 1B from Hugging Faceexport_llama32_qnn.py- Exports to ExecuTorch .pte format with QNN backendexport_llama32_qnn.bat- Windows automation script
app/src/main/cpp/executorch_llama32_qnn.cpp- Complete inference engine- Tokenizer, sampling, JNI bindings
- Updated CMakeLists.txt
ExecutorTorchLlama32.kt- Kotlin wrapper with coroutinesMainViewModel.kt- Multi-model state managementMainActivityEnhanced.kt- UI with model selection
The implementation includes placeholder code that will work without actual model files:
# 1. Build (running now...)
./gradlew assembleDebug
# 2. Install
adb install app/build/outputs/apk/debug/app-debug.apk
# 3. Test
# - Open app
# - Select "LLaMA 3.2 1B" from spinner
# - Enter a prompt: "What is machine learning?"
# - Tap "Generate"
# - You'll see simulated output (no real model needed)To test with the actual LLaMA 3.2 1B model:
# 1. Export model (requires ExecuTorch + QNN setup)
python export_llama32_qnn.py --model-dir models/llama-3.2-1b --quantize 16a4w
# 2. Push to device
adb shell mkdir -p /sdcard/EdgeAI/models/llama32
adb push output/llama32_qnn/llama32_1b_qnn.pte /sdcard/EdgeAI/models/llama32/
adb push output/llama32_qnn/tokenizer/tokenizer.model /sdcard/EdgeAI/models/llama32/
# 3. Grant permissions
adb shell pm grant com.example.edgeai android.permission.READ_EXTERNAL_STORAGE
adb shell pm grant com.example.edgeai android.permission.WRITE_EXTERNAL_STORAGE
# 4. Install and run
adb install app/build/outputs/apk/debug/app-debug.apk- App builds without errors
- Native library includes llama32 symbols
- No missing dependencies
- App launches successfully
- Model selection spinner works
- LLaMA 3.2 UI shows correctly
- Can enter prompts
- Generate button works (simulated or real)
- Response displays in TextView
- CLIP model still works
- Monitor logcat for initialization logs
- Check memory usage
- Verify inference speed (if using real model)
- No crashes or memory leaks
# Watch all EdgeAI logs
adb logcat | grep -E "(ExecutorTorchLlama32|MainViewModel|EdgeAI)"
# Watch LLaMA 32 specific logs
adb logcat | grep "ExecutorTorchLlama32"
# Clear logs and watch
adb logcat -c && adb logcat | grep "EdgeAI"I MainViewModel: 🚀 Initializing EdgeAI models...
I MainViewModel: Initializing LLaMA 3.2 1B...
I ExecutorTorchLlama32: 🚀 Initializing LLaMA 3.2 1B with ExecuTorch + QNN
I ExecutorTorchLlama32: Step 1: Loading tokenizer...
I ExecutorTorchLlama32: ✅ Tokenizer initialized (vocab_size=128256)
I ExecutorTorchLlama32: Step 2: Loading ExecuTorch model...
I ExecutorTorchLlama32: ✅ ExecuTorch model loaded
I ExecutorTorchLlama32: ✅ LLaMA 3.2 1B initialization complete!
I ExecutorTorchLlama32: 🎯 Generating response for prompt: "What is machine learning?"
I ExecutorTorchLlama32: 📝 Tokenized input: X tokens
I ExecutorTorchLlama32: ✅ Generation complete!
I ExecutorTorchLlama32: Generated X tokens in Yms (Z tokens/sec)
Issue: "Cannot find ExecutorTorchLlama32"
# Check file exists
ls app/src/main/java/com/example/edgeai/ml/ExecutorTorchLlama32.kt
# Clean and rebuild
./gradlew clean buildIssue: CMake errors
# Check CMakeLists.txt includes the new file
cat app/src/main/cpp/CMakeLists.txt | grep llama32Issue: "Model not initialized"
- Check if model files exist on device
- Verify permissions granted
- Check logcat for initialization errors
Issue: App crashes on launch
- Check logcat for stack trace
- Verify native library loaded correctly
-
Test Placeholder Mode First
- Verify UI works
- Check button interactions
- Confirm no crashes
-
Export Real Model
- Set up ExecuTorch environment
- Run export scripts
- Test with actual model
-
Performance Benchmarking
- Measure tokens/sec
- Monitor memory usage
- Test on different devices
-
UI Refinements
- Add streaming text display
- Improve progress indicators
- Add model switching
EdgeAI/
├── download_llama32_model.py ✅
├── export_llama32_qnn.py ✅
├── export_llama32_qnn.bat ✅
├── LLAMA32_SETUP.md ✅
├── app/src/main/
│ ├── cpp/
│ │ ├── executorch_llama32_qnn.cpp ✅
│ │ └── CMakeLists.txt ✅ (updated)
│ └── java/com/example/edgeai/
│ ├── ml/ExecutorTorchLlama32.kt ✅
│ ├── MainActivityEnhanced.kt ✅
│ └── ui/theme/MainViewModel.kt ✅ (updated)
- Full setup guide: LLAMA32_SETUP.md
- Implementation details: walkthrough.md
- Task checklist: task.md