Geeky Gadgets iconGeeky GadgetsSep 22, 2026 ~6 min source read

Old Apple Watch Series 6 Runs Falcon H1 LLM, Outperforms Raspberry Pi on Token Speed

Better Stack adapted llama.cpp to run the 90M-parameter Falcon H1 model on a 2020 Apple Watch Series 6, achieving 15–24 tokens per second by solving Core ML incompatibilities and tight memory limits.

Local AI on Old Apple Watch Series 6 Beats Raspberry Pi

Share this story

Send the public story page.

Useful takeaways from this story.

A six-year-old Apple Watch Series 6 ran the Falcon H1 model (90M parameters) at 15–24 tokens/sec after software and memory optimizations.

The Series 6 delivered token generation roughly 50× faster than a first-generation Raspberry Pi in this test, showing wearables can outpace single-board computers for some workloads.

# What happened Better Stack demonstrated that a Falcon H1 language model with 90 million parameters can run locally on an Apple Watch Series 6 (released 2020). To do that, they adapted the llama.cpp codebase to work around Core ML incompatibilities and the watch's constrained memory and architecture.

# Why this matters Running models on-device reduces reliance on cloud servers, cuts latency, and keeps data on the device. The experiment shows older wearables, when paired with lightweight models and targeted engineering, can perform useful inference tasks despite limited RAM and a 32-bit address space.

# How they did it The team tried to use Apple's Core ML but hit compatibility problems because Falcon H1's Mamba 2 state-space architecture isn't supported. Instead of forcing Core ML, they ported and rebuilt llama.cpp for watchOS. That required changing build flags, fixing architecture-specific issues, and squeezing memory usage to fit the watch's 1 GB of RAM and arm64_32 constraints.

They also tested a larger 135M-parameter model, but it proved less compatible with the device's resources. The smaller Falcon H1 was a better match for the watch's limits.

# Performance highlights

  • Token generation: 15–24 tokens per second on the Series 6 using Falcon H1.
  • Comparative result: token speed reported as about 50× faster than a first-generation Raspberry Pi in these specific tests.
  • Memory behavior: Falcon H1's memory consumption stayed consistent regardless of the number of tokens generated, which helped manage the watch's fixed RAM.

# Technical obstacles and fixes

  • Core ML incompatibility: Mamba 2 architecture could not be represented in Core ML. Solution: run the model via a modified llama.cpp build tailored to watchOS.
  • Memory addressing: The watch's arm64_32 mode and 1 GB RAM limit risked address-space exhaustion. Solution: aggressive memory optimization and selecting a smaller model that fits the addressable space.
  • Build and architecture issues: Required changing build flags and resolving platform-specific compile errors to produce a working binary for watchOS.

# Practical implications This experiment shows that on-device models can extend the usable life of older hardware by enabling new features without cloud access. For developers: choose model sizes that match device memory, expect to adapt existing inference runtimes for platform quirks, and measure token throughput versus user experience to decide if on-device inference is worthwhile.

# When this approach makes sense

  • Privacy-sensitive features where sending data to a server is unacceptable.
  • Situations where intermittent connectivity makes cloud reliance unreliable.

# Limits to expect

  • Larger models exceed addressable memory and will fail without quantization or other reductions.
  • Porting runtimes to constrained platforms takes engineering time and careful testing.

# Bottom line

More context around this story.

Apple Watch Ultra 4 получили новый чип впервые за два года
Computerra iconComputerraSep 9, 2026

Apple Watch Ultra 4 получили новый чип впервые за два года

Источник: Компьютерра - Журнал о науке и технологиях Apple Watch Ultra 4 получили новый процессор S11 — это первое серьезное обновление чипа в линейке за последние два года. Однако новый чип нужен Apple не только для того, чтобы сделать часы быстрее: он становится основой для целого набора новых интеллектуальных функци

Apple Unveils Watch Series 12 And Ultra 4 With S11 Chip
Ubergizmo iconUbergizmoSep 11, 2026

Apple Unveils Watch Series 12 And Ultra 4 With S11 Chip

Apple has officially introduced the Apple Watch Series 12 and Apple Watch Ultra 4 , bringing internal hardware upgrades, expanded artificial intelligence tools, and new casing options while maintaining the established design language of its wearable lineup.  Both models are powered by Apple’s next-generation S11 chi

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app