Multimodal AI
Visual& Voice
Hands-free, turn-by-turn voice conversations with the AI. Local Whisper transcription, 400+ neural voices across 80+ languages, push-to-talk barge-in, and voice transcripts saved as first-class session metadata. Plus paste-image visual analysis.

Continuous Voice Conversation Mode
Click the headset button in the chat input and just talk. Whisper transcribes locally, Plexon answers, the assistant’s reply is read aloud, and the mic auto-resumes, turn by turn, no buttons in between. Voice activity detection ends each turn the moment you pause; push-to-talk barge-in (Space by default) lets you cut the assistant off mid-reply and take over the conversation. Configure it once in Settings → Speech → Voice Mode, then run entire sessions hands-free.
- One click to start, hands-free turn-by-turn loop until you stop it
- Choose your TTS engine: Microsoft Edge (400+ neural voices, 80+ languages), My Voice (local voice cloning), or your OS’s native voices
- Auto-VAD ends each turn on silence; push-to-talk barge-in interrupts the assistant mid-reply
- Full voice transcripts saved as first-class session metadata (input_modality, language, model, duration), searchable and exportable
- Privacy-aware: raw audio sidecar is opt-in and off by default; only the transcript is kept

Local Speech-to-Text
Talk to your AI using Whisper, a state-of-the-art speech recognition model that runs entirely on your machine. No internet required after the initial model download. Choose from multiple model sizes to balance speed and accuracy.
- Fully offline. No audio data leaves your machine
- Multiple model sizes (base 142MB to large)
- Configurable language and initial prompt for domain accuracy
- Auto-send after transcription for hands-free operation

Text-to-Speech Output
Have the AI read responses aloud with high-quality neural voices. Choose from 400+ voices across 40+ languages, adjust speed, and optionally auto-read every new response.
- 400+ neural voices in 40+ languages
- Default engine + System Voices options
- Adjustable speed from 0.5x to 2x
- Auto-read new messages toggle
- Voice Activity Detection (VAD) for natural conversation flow

Visual Input: Images, Files & Screen Capture
Send visual context to the AI in multiple ways: paste from clipboard, drag & drop files onto the chat, click the photo icon to browse, or capture any area of your screen with Ctrl+Shift+S (Cmd+Shift+S on Mac). Supports JPEG, PNG, WebP, and GIF with up to 5 images per message.
- Screen capture: Ctrl+Shift+S to select and send any screen region
- Drag & drop: drop image files directly onto the chat input
- Clipboard paste: Ctrl+V to paste screenshots or copied images
- File picker: click the photo icon to browse and upload files
- Up to 5 images per message, auto-resized to optimize tokens
- Works with any vision-capable model across all providers

Image Processing & Optimization
Configure how images are processed before sending to the AI. Downscale large images (e.g. 4K → 1536px) to save ~6x tokens. Set compression quality for JPEG/WebP to balance quality and cost.
- Max dimension slider (512 to 4096px)
- Quality slider for JPEG/WebP compression
- PNG always sent lossless
- Significant token savings on high-resolution screenshots
More Multimodal Capabilities
Hands-Free Voice Loop
One click on the headset button starts a continuous turn-by-turn conversation. Speak, listen to the reply, the mic re-arms automatically. Stop the loop with the same button.
Push-to-Talk Barge-In
Hold Space (configurable) to interrupt the assistant mid-reply and start your next turn. Plexon stops the audio, captures your input, and resumes the loop without losing the thread.
Voice Transcripts as Session Metadata
Every spoken turn is tagged input_modality: voice with a voice_meta block (duration, model, language) and round-trips through session export. Raw audio sidecar is opt-in and off by default.
Screen Capture
Capture any area of your screen with Ctrl+Shift+S and send it directly to the AI for analysis. Perfect for debugging UI issues or explaining visual context.
Voice Activity Detection
VAD automatically stops recording when silence is detected. Combined with auto-send, this enables fully hands-free conversation with the AI.
80+ Languages
Microsoft Edge TTS covers 400+ neural voices across 80+ languages; Whisper STT recognises 90+. Pick a default voice and language per persona, or let the model auto-detect.