Multimodal AI
Visual& Voice
Hands-free, turn-by-turn voice conversations with the AI. Local Whisper transcription, 400+ neural voices across 80+ languages, push-to-talk barge-in, and voice transcripts saved as first-class session metadata. Plus paste-image visual analysis.
What happens in this clip
- The Speech tab holds voice chat, voice input and voice output in one place.
- A wake phrase starts voice mode by saying it, whether or not Plexon is in front.
- Five voice engines are listed with what each costs in download size or privacy.
- Plexon Voice is the cloud one, and it can clone your own voice from a short recording.
- Ask for a line to be recorded and the wav file is written to the project folder and plays back in the chat.
Related features

Continuous Voice Conversation Mode
Click the headset button in the chat input and just talk. Whisper transcribes locally, Plexon answers, the assistant’s reply is read aloud, and the mic auto-resumes, turn by turn, no buttons in between. Voice activity detection ends each turn the moment you pause; push-to-talk barge-in (Space by default) lets you cut the assistant off mid-reply and take over the conversation. Configure it once in Settings → Speech → Voice Mode, then run entire sessions hands-free.
- One click to start, hands-free turn-by-turn loop until you stop it
- Choose your TTS engine: Microsoft Edge (400+ neural voices, 80+ languages), Plexon Voice (Premium cloud voices, nothing to install), My Voice (local voice cloning), or your OS’s native voices
- Opening voice mode runs a quick readiness check and walks you through anything missing: mic access, the transcription model, your voices
- Optional hands-free activation: a system-wide shortcut, or say "Hey Plexon" from any app ("Bye Plexon" ends it). Both off by default
- Say "stop" or "pause" to cut a reply short, "start over" for a fresh session. No keyboard, no clicking
- Auto-VAD ends each turn on silence; push-to-talk barge-in interrupts the assistant mid-reply
- Speak commands, not just prompts: ask it to switch models or change a setting and it does it
- Full voice transcripts saved as first-class session metadata (input_modality, language, model, duration), searchable and exportable
- Privacy-aware: raw audio sidecar is opt-in and off by default; only the transcript is kept

Plexon Voice: Cloud Voices That Can Sound Like You
On Premium, Plexon Voice speaks with natural cloud voices and nothing to install: sign in and pick a voice. Record ten seconds of yourself and it clones your voice too. The clone shows up in the voice picker and works everywhere Plexon speaks, from read-aloud replies to persona voices. Free voice engines (neural, local cloning, system) stay available on every plan.
- Zero setup: no downloads, no GPU, works on any machine
- Curated natural voices, plus your own cloned voice from a short recording
- Works across read-aloud, auto-read, voice mode, and persona voices
- Ask by voice and it acts: "switch to the premium model" changes the setting instead of explaining it

Hands-Free, From "Hey Plexon" to "Bye Plexon"
Turn on the optional activation surfaces and voice mode starts without touching the keyboard: a system-wide shortcut and a wake phrase, both working from whatever app you are in. Say "Hey Plexon" and the window comes back with the conversation already running. Say "stop" to cut a reply short, "start over" for a clean session, "Bye Plexon" to end it. Every opening runs a quick readiness check that walks you through anything missing, so the first run works instead of failing quietly.
- Wake phrase and shortcut both work from any app, including when Plexon is minimised or in the tray
- The wake phrase holds the mic for as long as it is switched on, which is why it ships off
- "Stop" and "pause" cut a spoken reply short; "start over" opens a fresh session and keeps the old one
- "Bye Plexon" ends the conversation without sending it as a message
- Off by default; phrases, keys, and spoken commands are configurable in Settings → Speech

Local Speech-to-Text
Talk to your AI using Whisper, a state-of-the-art speech recognition model that runs entirely on your machine. No internet required after the initial model download. Choose from multiple model sizes to balance speed and accuracy.
- Fully offline. No audio data leaves your machine
- Eight models from 32MB to 1.08GB, listed fastest first with what each one trades away
- One model per speed rung for English and one for every other language, so a fast spoken turn is not an English-only privilege
- A large model for the languages the smaller ones struggle with, Greek among them
- Every model on the list beat the alternatives on a measured comparison, so a bigger download is never a slower, worse result
- Configurable language and initial prompt for domain accuracy
- Auto-send after transcription for hands-free operation

Text-to-Speech Output
Have the AI read responses aloud with high-quality neural voices. Choose from 400+ voices across 40+ languages, adjust speed, and optionally auto-read every new response.
- 400+ neural voices in 40+ languages
- Plexon Voice on Premium: natural cloud voices with zero setup, plus cloning your own voice from a short recording
- Ask Plexon to say something aloud in any chat and it speaks in the voice you picked, no voice mode needed
- Each engine remembers its own voice, so switching between them and back keeps your choice
- Default engine + System Voices options
- Adjustable speed from 0.5x to 2x
- Auto-read new messages toggle
- Voice Activity Detection (VAD) for natural conversation flow

Visual Input: Images, Files & Screen Capture
Send visual context to the AI in multiple ways: paste from clipboard, drag & drop files onto the chat, click the photo icon to browse, or capture any area of your screen with Ctrl+Shift+S. Supports JPEG, PNG, WebP, and GIF. Attach as many as fit in 24 MB per message, and the same picture twice counts once.
- Screen capture: Ctrl+Shift+S to select and send any screen region
- Drag & drop: drop image files directly onto the chat input
- Clipboard paste: Ctrl+V to paste screenshots or copied images
- File picker: click the photo icon to browse and upload files
- Up to 24 MB of pictures and clips per message, however many files that is
- Works with any vision-capable model across all providers

Image Processing & Optimization
Pictures go at their original size unless you turn downscaling on. Once it is on, large images are scaled to the longest side you pick (1920px by default), which cuts what a 4K screenshot costs. Compression quality for JPEG and WebP is a separate slider.
- Off by default: pictures are sent at their original size until you turn it on
- Max dimension slider (512 to 4096px, 1920 by default)
- Quality slider for JPEG/WebP compression
- PNG always sent lossless
- Significant token savings on high-resolution screenshots
More Multimodal Capabilities
Hands-Free Voice Loop
One click on the headset button starts a continuous turn-by-turn conversation. Speak, listen to the reply, the mic re-arms automatically. Stop the loop with the same button.
Push-to-Talk Barge-In
Hold Space (configurable) to interrupt the assistant mid-reply and start your next turn. Plexon stops the audio, captures your input, and resumes the loop without losing the thread.
Say "Hey Plexon"
Opt in to a wake phrase that starts voice mode from whatever app you are in and brings the window back, a system-wide shortcut that does the same on a keypress, and "Bye Plexon" to end the conversation. Once on, the wake phrase keeps listening, so it ships off.
Voice Transcripts as Session Metadata
Every spoken turn is tagged input_modality: voice with a voice_meta block (duration, model, language) and round-trips through session export. Raw audio sidecar is opt-in and off by default.
Screen Capture
Capture any area of your screen with Ctrl+Shift+S and send it directly to the AI for analysis. Perfect for debugging UI issues or explaining visual context.
Voice Activity Detection
VAD automatically stops recording when silence is detected. Combined with auto-send, this enables fully hands-free conversation with the AI.
80+ Languages
Microsoft Edge TTS covers 400+ neural voices across 80+ languages; Whisper STT recognises 90+. Pick a default voice and language per persona, or let the model auto-detect.
Connects to
Nothing here works alone. See the whole map for how the pieces fit together.
What feeds it
Customization feeds this.
The engine and the voice every spoken reply uses are set here.
Where it goes
This feeds AI Chat & Modes.
Speak instead of typing; local Whisper transcription becomes the message.
This feeds Video Composition.
Narration is spoken in whichever voice you set, including one cloned from your own recording.
This feeds Music Generation.
Clone your voice and ask for a song sung in it.