# What happened Better Stack demonstrated that a Falcon H1 language model with 90 million parameters can run locally on an Apple Watch Series 6 (released 2020). To do that, they adapted the llama.cpp codebase to work around Core ML incompatibilities and the watch's constrained memory and architecture.
# Why this matters Running models on-device reduces reliance on cloud servers, cuts latency, and keeps data on the device. The experiment shows older wearables, when paired with lightweight models and targeted engineering, can perform useful inference tasks despite limited RAM and a 32-bit address space.
# How they did it The team tried to use Apple's Core ML but hit compatibility problems because Falcon H1's Mamba 2 state-space architecture isn't supported. Instead of forcing Core ML, they ported and rebuilt llama.cpp for watchOS. That required changing build flags, fixing architecture-specific issues, and squeezing memory usage to fit the watch's 1 GB of RAM and arm64_32 constraints.
They also tested a larger 135M-parameter model, but it proved less compatible with the device's resources. The smaller Falcon H1 was a better match for the watch's limits.
# Performance highlights
- Token generation: 15–24 tokens per second on the Series 6 using Falcon H1.
- Comparative result: token speed reported as about 50× faster than a first-generation Raspberry Pi in these specific tests.
- Memory behavior: Falcon H1's memory consumption stayed consistent regardless of the number of tokens generated, which helped manage the watch's fixed RAM.
# Technical obstacles and fixes
- Core ML incompatibility: Mamba 2 architecture could not be represented in Core ML. Solution: run the model via a modified llama.cpp build tailored to watchOS.
- Memory addressing: The watch's arm64_32 mode and 1 GB RAM limit risked address-space exhaustion. Solution: aggressive memory optimization and selecting a smaller model that fits the addressable space.
- Build and architecture issues: Required changing build flags and resolving platform-specific compile errors to produce a working binary for watchOS.
# Practical implications This experiment shows that on-device models can extend the usable life of older hardware by enabling new features without cloud access. For developers: choose model sizes that match device memory, expect to adapt existing inference runtimes for platform quirks, and measure token throughput versus user experience to decide if on-device inference is worthwhile.
# When this approach makes sense
- Privacy-sensitive features where sending data to a server is unacceptable.
- Situations where intermittent connectivity makes cloud reliance unreliable.
# Limits to expect
- Larger models exceed addressable memory and will fail without quantization or other reductions.
- Porting runtimes to constrained platforms takes engineering time and careful testing.
# Bottom line