Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Tristam R.’s compact, 3D-printed smart speaker is a DIY Home Assistant voice satellite: an ESP32-based device that captures speech and plays responses while Home Assistant handles the heavier voice processing. The original build uses an ESP32-LyraT board; a later ESP32-S3 version adds a clearly documented local wake-word path. Neither is a self-contained Alexa replacement, and “local” depends on choosing local services for the full voice pipeline.
What Tristam built
The project packages a microphone-equipped ESP32 board, a small speaker driver, a status LED and a custom enclosure into a compact tabletop voice endpoint. The printed case has a speaker opening, microphone openings, an LED cutout and access for USB-C; the original design also leaves the board’s capacitive-touch controls accessible. A fabric grille covers the front. Tristam shared printable enclosure files through Printables; check the linked project page for the current files and license terms before making or sharing a derivative. Tristam’s original build guide describes the physical design and parts.
This is a maker project, not a commercial product or a tested plug-and-play kit. Its most useful distinction is that it exists in two hardware generations: an integrated ESP32-LyraT build and a later, more modular ESP32-S3 build.
Two designs, different trade-offs
| Choice | Original ESP32-LyraT | Later ESP32-S3 |
|---|---|---|
| Closest to Tristam’s first build | Yes | No; it is a separate design path |
| Audio hardware | Board includes audio circuitry, microphones and amplification | External I²S microphone and MAX98357A amplifier/DAC |
| Wiring | Less external audio wiring, but board-specific | More wiring and pin configuration |
| Wake-word detail | Do not assume the later local-wake-word setup applies | Example uses ESPHome micro_wake_word locally |
| Best fit | Faithful replication if the exact board is available | Adaptable prototype if you can verify an S3 board’s PSRAM and GPIOs |
The LyraT is not interchangeable with a generic ESP32 development board: its integrated audio hardware and pin assignments are central to that version. Tristam describes two 3-watt, 4-ohm outputs on the board, but verify the exact board revision and documentation before treating that as a purchasing specification. Original build details.
#1 Best Overall
- Please note!!! This product requires a 3.7V MX1.25 lithium battery for operation, which is not included. Please purchase it separately.
- High-Performance MCU: The board is equipped with the ESP32-S3R8 module, featuring a powerful Xtensa 32-bit LX7 dual-core processor that operates at up to 240MHz, ensuring efficient processing for various smart applications.
- Wireless Connectivity: With built-in support for 2.4GHz Wi-Fi (802.11 b/g/n) and Bluetooth 5 (LE), the ESP32-S3-AUDIO-Board offers robust wireless capabilities, facilitated by the onboard antenna for seamless communication and connectivity.
- Advanced Voice Interaction: The dual microphone array is designed with noise reduction and echo cancellation features, enabling accurate speech recognition and responsive near/far-field wake-up functionality, perfect for voice-activated applications.
- Dynamic Lighting Effects: Equipped with 7x programmable surround RGB LEDs, the board allows the creation of vibrant and colorful lighting effects, enhancing user interaction and visual appeal for projects.
The later project uses an ESP32-S3 development board, a MAX98357A I²S amplifier/DAC, an INMP441 I²S microphone, a Dayton Audio DMA45-4 1.5-inch, 4-ohm driver and a WS2812-compatible LED stick. This modular approach can be easier to source or adapt, but it makes correct power, wiring, microphone channel and GPIO configuration your responsibility. See Tristam’s later ESP32-S3 guide.
Where the voice processing happens
The ESP32 is the audio endpoint, not the main voice-computing machine. In broad terms, the path is:
Microphone → ESP32 → Home Assistant → speech recognition and conversation handling → speech synthesis → ESP32 speaker
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #2
- Compact and Portable: The ATOM VOICE is designed with a small form factor, measuring only 24 * 24 * 17 mm. Its compact size makes it highly portable and convenient for on-the-go use.
- Voice Interaction and AI Capabilities: The built-in microphone and speaker allow for voice interaction, enabling voice control, story-telling, and other AI-based functions. The device can be programmed to access cloud platforms like AWS and Baidu, expanding its capabilities.
- Wireless Music Playback: Utilizing the BT capabilities of the ESP32, you can wirelessly play music from your mobile phone or tablet, providing a seamless and convenient audio experience.
- Versatile Connectivity: The ATOM VOICE supports 2.4G Wi-Fi IEEE 802.11b/g/n, allowing for easy and reliable wireless connectivity to the internet and other devices.
- RGB LED Status Display: The embedded RGB LED (SK6812) visually displays the connection status, providing a clear indication of the device's operational mode and status.
- On the satellite: capture microphone audio, connect over Wi-Fi, play audio returned by Home Assistant, and signal status with LEDs. In the later S3 configuration,
micro_wake_worddetects the configured wake word on the device. - On the Home Assistant host: receive the request, run the configured speech-to-text service, handle the command or conversation, generate speech and return it to the satellite. The original project description identifies Whisper for speech-to-text and Piper for text-to-speech on the Home Assistant server—not on the ESP32. Hackster’s overview of the original project.
“Local” is conditional, not a blanket privacy guarantee. With local speech-to-text, local text-to-speech and a local conversation agent, the voice pipeline can stay on your home network rather than relying on Alexa- or Google-operated speech services. Choosing a cloud agent or external service changes that. The satellite still needs Wi-Fi and a working Home Assistant server; integrations or remote access may also communicate externally.
Local wake-word detection can prevent continuously sending microphone audio to the server just to decide whether you are addressing the device. It does not make an always-powered microphone risk-free: network security, firmware, logs and physical access still matter. The wake word is only the activation stage; it does not perform transcription, understand the request or synthesize a reply.
Parts and enclosure considerations
| Part | Role and caveat |
|---|---|
| ESP32-LyraT board (original) | Controller and integrated audio platform; confirm revision, microphones, amplifier and pinout. |
| ESP32-S3 board (later) | Controller; select a documented board with suitable PSRAM and accessible GPIOs rather than assuming any S3 matches the example. |
| INMP441 microphone (later) | I²S audio input; channel selection, voltage and breakout pin labels matter. |
| MAX98357A (later) | I²S amplifier/DAC; match supply, wiring and speaker impedance. |
| Dayton Audio DMA45-4 | Small 1.5-inch, 4-ohm full-range driver used in both descriptions; a small driver is suited to voice responses, not necessarily loud-room use or music with strong bass. |
| Addressable LED stick | Status indication for states such as listening, mute, error and completion; LED count, data protocol and power draw must match the setup. |
| Printed enclosure, grille and fasteners | Frame and speaker housing, fabric, adhesive and mechanical fasteners; test-fit parts before permanent assembly. |
Enclosure openings affect both sound and microphone pickup. Keep the microphone ports clear, avoid placing the speaker so it feeds directly into the microphones, and leave USB access available for setup and recovery. Test the electronics before adding fabric or adhesive; padding can alter acoustics and should be added only after the basic audio path works.
Rank #3
- ESP32-S3-AUDIO-Board adopts ESP32-S3R8 module with 32-bit LX7 dual-core processor, up to 240MHz main frequency. Supports 2.4GHz Wi-Fi (802.11 b/g/n) and Bluetooth 5 (LE), with onboard antenna
- Integrated 512KB Static RAM, 384KB ROM, 8MB PSRAM, and external 16MB Flash memory. Onboard TF card slot for storing audio files, etc.
- Onboard Dual microphone array with noise reduction and echo cancellation, suitable for accurate speech recognition and near/far-field wake-up. Onboard audio decoding chip, dual microphones and speaker header. Onboard 7x surround RGB LEDs, programmable for a variety of dynamic effects
- Onboard SPI LCD display interface (FPC connector / pin header), DVP camera interface (24pin connector), USB, I2C, and some I/O pins (compatible with display interface I/O pins). Onboard multiple reserved buttons and battery switch for customized function development
- Integrated PCF85063 RTC chip, supports power-off time retention for alarm, scheduled task, and wake-up functions. Built-in battery recharge management module, supports multiple power modes and low-power applications
Build it in dependency order
- Choose a generation. Pick the LyraT for the closest reproduction, provided you have the right board revision. Pick the ESP32-S3 path for a modular build and local wake-word example, accepting additional wiring.
- Verify the exact board. Check its pinout, PSRAM, audio connections and power requirements. Marketplace listings can use the same board name for different revisions or configurations.
- Make Home Assistant work first. Configure an Assist pipeline. If you want cloud-independent speech, configure local STT, TTS and conversation handling and confirm the host can run the selected models. The ESP32 does not run Whisper or Piper.
- Wire the audio and LEDs. Follow the actual board pinout, not a copied pin list. On the S3 version, connect the I²S microphone and MAX98357A, observe speaker polarity and impedance, and provide suitable power and a common ground. Wire the LED data and power for the strip you actually use.
- Adapt the ESPHome configuration. Start from Tristam’s S3 example or the original project’s configuration, then verify it against your board and installed ESPHome release. Replace Wi-Fi credentials and API encryption key; check the target, framework, PSRAM mode, pins and audio settings. Do not treat an example YAML as universal or guaranteed current.
- Flash by USB, then adopt the device. Use serial logs during initial setup. Confirm the device joins Wi-Fi and Home Assistant can see it before closing the case. Keep USB recovery possible even after OTA is enabled.
- Test each stage separately. Check boot, Wi-Fi, Home Assistant adoption, LEDs, speaker playback, microphone capture, wake word, transcription, a simple command and spoken reply—in that order. This isolates hardware failures from pipeline failures.
- Finish the enclosure. Test-fit the printed parts, confirm mic openings and USB access are unobstructed, align the driver and LED, then attach the grille and fasten the case.
- Secure the endpoint. Use encrypted Home Assistant API communication, protect secrets and OTA access, and disable optional interfaces such as a web server when they are no longer needed. Put the device on a trusted network or appropriately isolated IoT network.
What the later example configuration tells you
Tristam’s S3 example is a useful map of the design, not a universal recipe. It specifies an ESP32-S3 using ESP-IDF, a 240 MHz CPU setting and octal PSRAM at 80 MHz; encrypted Home Assistant API access; Wi-Fi secrets; OTA, captive portal and optional web server support; and an eight-LED WS2812 strip on GPIO16. Its I²S microphone uses GPIO6 for word select, GPIO7 for bit clock and GPIO4 for data input; speaker output to the MAX98357A uses GPIO8. The example sets micro_wake_word to hey_jarvis, noise suppression to 2.0 and voice volume multiplier to 4.0, and changes LED states for wake, mute, error and completion.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteThose values belong to his chosen hardware and published configuration. GPIO numbers do not transfer across boards, and framework defaults, PSRAM support, audio components and YAML syntax can change between ESPHome releases. Check the current documentation and the exact board before flashing. The project’s June 2025 configuration update is not proof that the same file will work unchanged with every 2026 board or software version.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting by symptom
The device does not appear in Home Assistant
- Confirm 2.4 GHz Wi-Fi credentials and that the board received an IP address.
- Check the API encryption key and serial logs over USB.
- Verify the firmware target and board definition; a pin or PSRAM mismatch can prevent reliable boot.
The wake word triggers, but no command is handled
First run the Assist pipeline from Home Assistant itself. If that fails, fix the server-side pipeline before debugging the satellite. Then check that wake detection starts the voice-assistant component, the selected STT service is available, the microphone format and channel are correct, and the chosen conversation agent can process the request.
Rank #4
- POWERFUL PROCESSOR: Integrated ESP32-S3-PICO-1-N8R8 dual-core Xtensa LX7 processor running up to 240MHz with 8MB Flash and 8MB PSRAM for robust IoT voice interaction applications
- HIGH-QUALITY AUDIO: ES8311 24-bit mono audio codec with NS4150B Class D amplifier and 8 1W speaker delivers clear voice pickup through MEMS microphone (65dB SNR) and high-fidelity audio output
- WIRELESS CONNECTIVITY: Built-in 2.4GHz Wi-Fi enables seamless wireless communication for smart home control, AI voice assistants, and human-computer interaction scenarios
- COMPACT DESIGN: Ultra-compact dimensions of 0.94 x 0.94 x 0.66 inches and lightweight at 0.33 ounces, with integrated infrared transmitter for enhanced control capabilities
- EXPANDABLE INTERFACE: Programmable IoT controller with expandable pins and interfaces, operating at DC 5V power input and temperature range of 32F to 104F for versatile development needs
There is static, hum or weak playback
Tristam’s later guide notes static reports and says to try powering the MAX98357A at 5 V rather than the 3.3 V used in an earlier attempt; the issue was not consistently reproducible. Treat this as a troubleshooting lead, not a universal fix, and respect the specific amplifier breakout’s ratings. Later build notes.
- Check common ground, adequate USB power, short audio/power wiring and the configured I²S pins.
- Separate LED power wiring from audio where practical; LEDs can introduce electrical noise.
- Confirm speaker impedance and polarity, and reduce volume to rule out clipping or overdriving the small driver.
Wake-word recognition is unreliable
Check mic orientation, clear enclosure openings, microphone channel selection, gain, distance and background noise. Speaker feedback or noisy LED/amplifier power can also hurt recognition. Try a suitable model and wake word; detection is not guaranteed in every room or at every distance.
An OTA update fails
OTA depends on the device staying on the network and accepting the new firmware. Keep a USB cable and serial-flashing route available: a bad configuration or lost network connection may require reflashing over USB.
Best Value
- Adopts ESP32-S3R8 module with Xtensa 32-bit LX7 dual-core processor, up to 240MHz main frequency. Supports 2.4GHz Wi-Fi (802.11 b/g/n) and Bluetooth 5 (LE), with onboard antenna. Integrated 512KB SRAM, 384KB ROM, 8MB PSRAM, and external 16MB Flash memory.
- AI Voice Interaction: Dual microphone array with noise reduction and echo cancellation, suitable for accurate speech recognition and near/far-field wake-up. Supports AI Speech Interaction: Allows access to online large model platforms such as DeepSeek, GPT, Doubao, etc
- Onboard Audio Input/Output: Supports high-quality audio processing, providing clear and high-quality audio input and output. Equipped with the offline voice model we provided to realize device control via customizable shortcut commands.
- Colorful Lighting Effects: Onboard 7x surround RGB LEDs, programmable for a variety of dynamic effects. Clock Management: Integrated PCF85063 RTC chip, supports power-off time retention for alarm, scheduled task, and wake-up functions. HMI Interfaces: Multiple reserved buttons and battery switch for customized function development.
- Supports External LCD Displays & Cameras: Onboard LCD interface, compatible with Wave-share 1.47inch / 2inch / 2.8inch / 3.5inch LCDs and other SPI displays. Onboard DVP interface, compatible with ESP32 OV2640 / OV5640 cameras.
Is this a good build for you?
It is a compelling project if you already use Home Assistant, want to learn ESPHome audio, value local control and are comfortable with board-specific wiring and YAML. The later S3 build is the more adaptable starting point when you can verify a suitable board, but it is not simpler in every respect. The LyraT can reduce external audio wiring, yet its integrated design makes an exact board match important.
Choose a ready-made Home Assistant voice device if you want a faster setup, predictable support and less fabrication risk. Choose a commercial smart speaker if acoustics and effortless setup matter more than an open, locally configured design—but understand its account and cloud trade-offs. This small DIY speaker should not be expected to match commercial devices’ microphone pickup, volume, bass response or reliability. Historical estimates in the project posts—including the creator’s under-US$50 estimate for the later build—are not current prices; availability, regional costs and compatible board revisions need checking before buying.
Useful starting points: Home Assistant voice control, Home Assistant, ESPHome, the original LyraT guide and the later ESP32-S3 guide.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.


Leave a Reply