Audio & Podcast Transcriber
In the Apify Store
Audio files, video files and podcast feeds to text: transcript, timed segments, SRT and WebVTT, from open-source Whisper models with a cap on minutes.
Give it direct links to audio or video files, or a podcast RSS feed and the number of latest episodes. It returns one row per file or episode: the transcript, timed segments, ready SRT and WebVTT subtitles, and the detected language. It runs open-source Whisper models on Apify in two sizes, base (fast, cheapest) and small (slower, more accurate). No third-party transcription API is called.
What you get
- Transcript as one text field per file or episode, with the detected language and Whisper's confidence in it, or the language you set.
- Timed segments: a list of
{start, end, text}with times in seconds, for search, quotes and chapter marks. - SRT and WebVTT subtitle files as text, ready to save next to a video.
- Podcast feeds: the latest 1 to 50 episodes of each RSS or Atom feed, with title, guid, publication date and feed link on every row. With New episodes only on, a scheduled run transcribes just the episodes that appeared since the last run.
- Two models:
baseandsmall, and an optional translation to English with Whisper's built-in translation task. - A hard cap on minutes: Max total minutes (default 120) limits the audio you pay for in a run. A file that would pass it is cut at the cap, and the files after it get a free
budget-reachederror row. - Error rows for anything that fails (blocked link, 404, not audio, robots.txt) with the reason and what to do. They are free and never stop the run.
Real output
{
"recordType": "transcript",
"status": "ok",
"sourceIndex": 0,
"sourceUrl": "https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg",
"feedUrl": null,
"episodeTitle": null,
"episodeGuid": null,
"publishedAt": null,
"fileSizeBytes": 1139395,
"durationSec": 133.2,
"language": "en",
"languageProbability": 0.993,
"model": "base",
"task": "transcribe",
"text": "This is a Libravox Recording. All Libravox recordings are in the public domain. For more information or to volunteer, please visit Libravox.org. This reading by Michael Sherrer, ...",
"segments": [
{
"start": 1.62,
"end": 7.46,
"text": "This is a Libravox Recording. All Libravox recordings are in the public domain."
},
{
"start": 7.46,
"end": 12.46,
"text": "For more information or to volunteer, please visit Libravox.org."
},
{
"start": 12.46,
"end": 19.78,
"text": "This reading by Michael Sherrer, www.americanafonic.com"
}
],
"srt": null,
"vtt": null,
"wordCount": 302,
"billedMinutes": 3,
"processingSec": 6.5,
"warnings": [],
"error": null,
"processedAt": "2026-10-03T10:43:12Z"
}
{
"recordType": "transcript",
"status": "ok",
"sourceIndex": 6,
"sourceUrl": "https://www.archive.org/download/round_moon_librivox/roundthemoon_00_verne_64kb.mp3",
"feedUrl": "https://librivox.org/rss/1078",
"episodeTitle": "Preliminary Chapter",
"episodeGuid": null,
"publishedAt": null,
"fileSizeBytes": 4917416,
"durationSec": 120,
"language": "en",
"languageProbability": 0.993,
"model": "base",
"task": "transcribe",
"text": "Preliminary chapter, Round the Moon. This is a Libravox recording. All Libravox recordings are in the public domain. For more information or to volunteer, please visit Libravox.org recorded by Larry ...",
"segments": [
{
"start": 0.98,
"end": 4.02,
"text": "Preliminary chapter, Round the Moon."
},
{
"start": 4.02,
"end": 6.02,
"text": "This is a Libravox recording."
},
{
"start": 6.02,
"end": 8.78,
"text": "All Libravox recordings are in the public domain."
}
],
"srt": "1\n00:00:00,980 --> 00:00:04,020\nPreliminary chapter, Round the Moon.\n\n2\n00:00:04,020 --> 00:00:06,020\nThis is a Libravox recording.\n\n3\n00:00:06,020 --> 00:00:08,780\nAll Libravox recordings are in the public\ndomain.\n\n...",
"vtt": "WEBVTT\n\n00:00:00.980 --> 00:00:04.020\nPreliminary chapter, Round the Moon.\n\n00:00:04.020 --> 00:00:06.020\nThis is a Libravox recording.\n\n00:00:06.020 --> 00:00:08.780\nAll Libravox recordings are in the public\ndomain.\n\n00:00:08.780 --> 00:00:13.380\nFor more information or to volunteer,\nplease visit Libravox.org\n\n00:00:13.380 --> 00:00:15.980\nrecorded by Larry Ann Walden.\n\n00:00:15.980 --> 00:00:18.980\nRound the Moon by Jules Verne.\n\n00:00:18.980 --> 00:00:25.980\nPreliminary chapter, recapitulating the\nfirst part of this work and serving as a\npreface to the second.\n\n00:00:26.980 --> 00:00:34.980\nDuring the year 1860, blank, the whole\nworld was greatly excited by a scientific\nexperiment unprecedented in the annals of\nscience.\n\n00:00:34.980 --> 00:00:48.980\nThe members of the gun club, a circle of\nartillery men formed at Baltimore after\nthe American War, conceived the idea of\nputting themselves in communication with\nthe Moon, yes with the Moon, by sending to\nher a projectile.\n\n00:00:48.980 --> 00:01:05.980\nTheir president, Barbacaine, the promoter\nof the enterprise, having consulted the\nastronomers of the Cambridge Observatory\nupon the subject, took all necessary means\nto ensure the success of this\nextraordinary enterprise, which had been\ndeclared practicable by the majority of\ncompetent judges.\n\n00:01:05.980 --> 00:01:13.980\nAfter setting on foot a public\nsubscription, which realized nearly 1.2\nmillion pounds, they began the gigantic\nwork.\n\n00:01:13.980 --> 00:01:32.980\nAccording to the advice forwarded from the\nmembers of the Observatory, the gun\ndestined to launch the projectile had to\nbe fixed in a country situated between the\n0 and 28th degrees of North or South\nlatitude, in order to aim at the Moon when\nat the zenith, and its initiatory velocity\nwas fixed at 12,000 yards to the second.\n\n00:01:32.980 --> 00:01:59.980\nLaunched on the 1st of December at 10\nhours 46 minutes 40 seconds PM, it ought\nto reach the Moon four days after its\ndeparture, that is, on the 5th of\nDecember, at midnight precisely at the\nmoment of her attaining her paragy, that\nis, her nearest distance from the Earth,\nwhich is exactly 86,410 leagues French, or\n238,833 miles, mean distance, England.\n",
"wordCount": 298,
"billedMinutes": 2,
"processingSec": 20.8,
"warnings": [
"Only the first 2 minutes were transcribed (limit: 'Max minutes per file', 'Max total minutes' or the run's maximum cost)."
],
"error": null,
"processedAt": "2026-10-03T10:45:12Z"
}
{
"recordType": "error",
"status": "error",
"sourceIndex": 4,
"sourceUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"feedUrl": null,
"episodeTitle": null,
"episodeGuid": null,
"publishedAt": null,
"errorCode": "unsupported-source",
"error": "Links on youtube.com are not supported: Tinlark transcribes direct audio or video files and podcast RSS feeds, not YouTube or social-media pages.",
"processedAt": "2026-10-03T10:44:35Z"
}
Example input
Paste this into the input form, or send it through the API.
{
"podcastFeedUrls": ["https://librivox.org/rss/1078"],
"episodesPerFeed": 2,
"model": "small",
"language": "auto",
"outputFormats": ["text", "segments", "srt"],
"maxTotalMinutes": 60
}
Run it from code
One request to the Apify API starts a run and returns the dataset rows when it finishes. Use your own Apify API token. The Apify client libraries and Apify's MCP server work too.
API=https://api.apify.com/v2/acts/tinlark~audio-podcast-transcriber
curl -X POST "$API/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/3/30/LibriVox_-_Everrett_Copy_of_the_Gettysburg_Address_-_Michael_Scherer.ogg"], "model": "base", "outputFormats": ["text", "segments"]}'
Limits
- Direct links to media files and podcast feeds only. YouTube, Instagram, TikTok, Facebook and X links are not supported and give an
unsupported-sourceerror row, also when a link redirects there. - Files up to 1 GB, at most 600 minutes per file (180 by default). Longer audio is cut with a warning. For files over 180 minutes, raise the memory (up to 8192 MB).
- Public files only: no login, no cookies, no signed-in pages. Private network addresses are refused. One download per host at a time.
- No accuracy figure. The models are open-source Whisper; Tinlark has not measured word error rates, so none are claimed. Names and unusual words can be misspelled, more so with the smaller
basemodel. Noisy, accented or overlapping speech is harder;smalldoes better there thanbase. - Speech as spoken: no speaker names, no punctuation guarantees, no word timings. Long recordings can contain repeated or invented passages where the speech is unclear: check the output when it matters.
- Music-only or silent files return an empty transcript and a warning ("No speech was detected").
- Speed: on a 13 min 47 s MP3 the base model ran about 9 times faster than real time and small about 2.7 times, measured on Apify at 4096 MB (the default). Start-up takes about 15 to 20 seconds per run.
- New episodes only looks at the latest Episodes per feed episodes only. If a feed publishes more than that between two runs, raise the number.
Data source
Only the files and feeds you give it. It does not search, crawl or log in anywhere. It checks a feed host's robots.txt before fetching the feed, fetches media files exactly as linked, and identifies itself with a User-Agent that names Tinlark and gives a contact address. Files are downloaded into the run, transcribed there and deleted when it ends. Process only recordings you have the right to use: you are responsible for lawful use of the audio and the output.
Questions before you start?
Write to [email protected]. Each product page lists what the tool does, what it does not do, and its exact price.