# What happened Developer cpldcpu squeezed an image generation model onto the Raspberry Pi Pico 2's RP2350 microcontroller. The project, called Pico‑Faces, generates 128×128 RGB images of human faces and outputs them to a VGA monitor or streams them over USB.
# What the model is and how small it is Pico‑Faces implements a latent flow diffusion transformer (DiT), the same general architecture used in some larger image generation models but heavily compressed. Two released variants are 2.9 million and 1.7 million parameters. For context, typical local diffusion models are roughly 5,000 times larger.
# Performance and output
- Image size: 128×128 RGB.
- Generation speed: roughly 5–20 seconds per image on the Pico 2, depending on variant and settings.
- Conditional controls: supports five classes covering gender, smile, and neutral expressions.
Outputs can be sent to a VGA display using a Pico VGA adapter or streamed as image data over USB serial to a host machine for rendering with Python libraries.
# How to try it quickly
- 1Put the Pico 2 into BOOTSEL mode and copy the provided UF2 file to the RPI‑RP2 drive.
- 2On the desktop, install pyserial, matplotlib, numpy, and Pillow to run the viewer that renders images streamed over USB.
That lets you test the prebuilt firmware without rebuilding the model or needing a GPU.
# Rebuilding and retraining options
- Full retrain pipeline: dataset download, VAE and decoder training, latent generation, DiT training, calibration, and quantization‑aware training. The full retrain takes about a day on a CUDA GPU.
The developer notes that many optimizations used for larger models were also necessary to make the microcontroller build work.
# Why this matters for hobbyists and makers This project shows how model architecture choices plus compression and careful optimization can fit generative image models onto very constrained hardware. It opens practical possibilities for offline image generation in embedded projects where size, cost, or connectivity are limiting factors.
# Where to learn more Source code, checkpoints, and build instructions are available on the project's GitHub. The Adafruit post summarizes the project and links to the repository, and an accompanying Learn Guide expands on running Pico‑Faces on alternative video output hardware like Fruit Jam or HSTX/DVI.
# Practical uses and experiments to try
- Hook the Pico 2 to a small VGA display for a standalone generative art toy.
- Stream images to a PC and combine them with simple UI controls to change conditioning (gender, smile, neutral).
- Use the checkpoints as a starting point for further quantization or retraining targeted at custom datasets.
Pico‑Faces provides a tested firmware workflow and a retrain path for makers who want to explore tiny, on‑device generative models.