Launching first on macOS · Windows to follow

PipScribe

Bring every word back to the moment it was spoken.

A local AI transcription and subtitle workspace that turns long audio and video into searchable, editable, export-ready timed text.

See pricing & trial
Interview_July.mov Completed

Interview_July

47:18 · Qwen3-ASR · 2 speakers detected
TranscriptAI SummaryAI Chat
41
00:12:06 → 00:12:10Maya Reed

The timing is what makes the transcript useful later.

42
00:12:10 → 00:12:15Daniel Hart

Everywordstaysconnectedtothemomentitwasspoken.

43
00:12:15 → 00:12:19Maya Reed

That makes review much faster than scrubbing through the whole recording.

12:12 / 47:18
Local transcriptionSpeaker recognitionWord-level timestampsTranslation and subtitlesAI summaries and chat

Turn speech into working material

Not just text.Keep working.

From interviews to video footage, PipScribe keeps audio, text, and time connected so every later edit has a source.

Interviews and podcasts

Separate speakers and replay the original audio by segment.

Video and subtitles

Adjust timing, control subtitle length, and export.

Courses and research

Search long recordings and extract summaries and key ideas.

Meetings and archives

Batch transcribe, translate, and continue working with the content.

Bring almost anything

Start with a fileor a video URL.

Import local audio or video, or paste a supported video URL. PipScribe prepares the media and adds the transcription to your local queue.

Drop audio or video here

Supports common formats readable by FFmpeg

MP3WAVM4AMP4MOVMKVWEBMMore

Transcribe locally after retrieving the media

Precise timing

Hear the sentence.See the exact word.

Qwen3-ASR word-level timestamps keep playback, review, and subtitle editing on the same timeline.

Reviewtheexactwordwithoutlosingyourplace.

00:16:41.240
Possible long pause 00:16:42 → 00:16:47 gap 4.9s
00:16:42.000

After transcription

Know who spoke,shape the final subtitle.

Speaker labels, translations, and subtitle versions share one timeline. Continue the work in one place without moving content between tools.

VOICE MAP 2 speakers detected
Maya Reed

First, verify the word-level timeline. Then refine the subtitle boundaries.

Daniel Hart

That way, every edit can be checked against the original audio.

Rename and manage speaker labels for each project

Local models and our inference engine

Choose the model.Use the right hardware.

Match Parakeet, Qwen3, or Whisper to the task, then let PipScribe's own engine select the right local compute backend for your device.

Built for speed

Move through long recordings quickly. On an M3 MacBook, a 119-second test recording completed in as little as 3.3 seconds.
About 36× realtime Measured on an M3 MacBook Ideal for long recordings
LOCAL INFERENCE CORE PipScribe Engine

Handles model loading, audio processing, and inference scheduling, with CPU fallback when no compatible GPU is available.

APPLE SILICONNow
Metal

Accelerated on the Mac GPU.

NVIDIAWindows
CUDA

A dedicated NVIDIA backend.

AMD · INTELFlexible
Vulkan

Accelerated on supported GPUs.

The current macOS release uses Metal. CUDA and Vulkan support will arrive with the Windows release; availability depends on the GPU and driver environment.

Understand, organize, and ask more

Keep workingwith the transcript.

Summarize, extract key ideas, identify open questions, or chat with the current transcript. Edit the built-in prompts or create your own.

Interview_July · Current transcriptOllama · Local

Turn this interview into an actionable summary.

AI response

Every conclusion can return to the original audio.

Use word-level timing to locate the exact moment first, then review long pauses and subtitle boundaries. Summaries, corrections, and later edits all keep the same audio reference.

00:12:06 · Maya Reed: “First, verify the word-level timeline. Then refine the subtitle boundaries.”
Cloud or local. You decide.Use your own account or API key, or connect a model running on your computer.
OpenAIAnthropicGeminiQwenOpenRouterDeepSeekZ.aiOllama LocalLM Studio Local

A clear privacy boundary

Your original media stays on your computer.

Local by default. Online services connect only when you choose them.

Media is not uploadedAudio and video are processed on your computer
Local work stays localQwen3, Parakeet, Whisper, Ollama, and LM Studio
Online only by choiceTranslation and cloud AI run only when you enable them

Simple pricing

Buy once.Use every file.

Try every feature on your own files before deciding to buy.

Lifetime licenseUS$59
  • 7-day full-feature trial
  • Three families of local transcription models
  • Batch processing and word-level timestamps
  • Subtitles, translation, summaries, and AI chat

Online AI services use your own account or API key, and provider charges may apply. Local Ollama and LM Studio do not require a cloud API.

Frequently asked questions

Which audio and video formats are supported?

PipScribe supports common formats readable by FFmpeg, including MP3, WAV, M4A, MP4, MOV, MKV, and WebM. Available codecs may vary by platform.

Can I transcribe a video URL directly?

You can paste supported video URLs into the queue. Availability may change with the source website's rules; local file imports are unaffected.

Can I transcribe without an internet connection?

After downloading a model, local file transcription does not require uploading your media to the cloud.

Can AI summaries and translation run entirely locally?

Yes. Use Ollama or LM Studio installed on your computer, or choose an online service when needed.

When will the Windows version be available?

The Windows version is in development and will have its own download when it launches.