Multimodal AI

Visual& Voice

Hands-free, turn-by-turn voice conversations with the AI. Local Whisper transcription, 400+ neural voices across 80+ languages, push-to-talk barge-in, and voice transcripts saved as first-class session metadata. Plus paste-image visual analysis.

Voice Mode settings on the Speech tab

Continuous Voice Conversation Mode

Click the headset button in the chat input and just talk. Whisper transcribes locally, Plexon answers, the assistant’s reply is read aloud, and the mic auto-resumes, turn by turn, no buttons in between. Voice activity detection ends each turn the moment you pause; push-to-talk barge-in (Space by default) lets you cut the assistant off mid-reply and take over the conversation. Configure it once in Settings → Speech → Voice Mode, then run entire sessions hands-free.

  • One click to start, hands-free turn-by-turn loop until you stop it
  • Choose your TTS engine: Microsoft Edge (400+ neural voices, 80+ languages), My Voice (local voice cloning), or your OS’s native voices
  • Auto-VAD ends each turn on silence; push-to-talk barge-in interrupts the assistant mid-reply
  • Full voice transcripts saved as first-class session metadata (input_modality, language, model, duration), searchable and exportable
  • Privacy-aware: raw audio sidecar is opt-in and off by default; only the transcript is kept
Speech to Text settings on the Speech tab

Local Speech-to-Text

Talk to your AI using Whisper, a state-of-the-art speech recognition model that runs entirely on your machine. No internet required after the initial model download. Choose from multiple model sizes to balance speed and accuracy.

  • Fully offline. No audio data leaves your machine
  • Multiple model sizes (base 142MB to large)
  • Configurable language and initial prompt for domain accuracy
  • Auto-send after transcription for hands-free operation
Text to Speech settings on the Speech tab

Text-to-Speech Output

Have the AI read responses aloud with high-quality neural voices. Choose from 400+ voices across 40+ languages, adjust speed, and optionally auto-read every new response.

  • 400+ neural voices in 40+ languages
  • Default engine + System Voices options
  • Adjustable speed from 0.5x to 2x
  • Auto-read new messages toggle
  • Voice Activity Detection (VAD) for natural conversation flow
A chat with an image attached to the message box

Visual Input: Images, Files & Screen Capture

Send visual context to the AI in multiple ways: paste from clipboard, drag & drop files onto the chat, click the photo icon to browse, or capture any area of your screen with Ctrl+Shift+S (Cmd+Shift+S on Mac). Supports JPEG, PNG, WebP, and GIF with up to 5 images per message.

  • Screen capture: Ctrl+Shift+S to select and send any screen region
  • Drag & drop: drop image files directly onto the chat input
  • Clipboard paste: Ctrl+V to paste screenshots or copied images
  • File picker: click the photo icon to browse and upload files
  • Up to 5 images per message, auto-resized to optimize tokens
  • Works with any vision-capable model across all providers
Image Processing settings on the Integrations tab

Image Processing & Optimization

Configure how images are processed before sending to the AI. Downscale large images (e.g. 4K → 1536px) to save ~6x tokens. Set compression quality for JPEG/WebP to balance quality and cost.

  • Max dimension slider (512 to 4096px)
  • Quality slider for JPEG/WebP compression
  • PNG always sent lossless
  • Significant token savings on high-resolution screenshots

More Multimodal Capabilities

Hands-Free Voice Loop

One click on the headset button starts a continuous turn-by-turn conversation. Speak, listen to the reply, the mic re-arms automatically. Stop the loop with the same button.

Push-to-Talk Barge-In

Hold Space (configurable) to interrupt the assistant mid-reply and start your next turn. Plexon stops the audio, captures your input, and resumes the loop without losing the thread.

Voice Transcripts as Session Metadata

Every spoken turn is tagged input_modality: voice with a voice_meta block (duration, model, language) and round-trips through session export. Raw audio sidecar is opt-in and off by default.

Screen Capture

Capture any area of your screen with Ctrl+Shift+S and send it directly to the AI for analysis. Perfect for debugging UI issues or explaining visual context.

Voice Activity Detection

VAD automatically stops recording when silence is detected. Combined with auto-send, this enables fully hands-free conversation with the AI.

80+ Languages

Microsoft Edge TTS covers 400+ neural voices across 80+ languages; Whisper STT recognises 90+. Pick a default voice and language per persona, or let the model auto-detect.