About this project

# podgenai podgenai is a Python application that generates informational single-speaker or two-speaker audiobook/podcast MP3 files on a given topic using an OpenAI LLM. Content comes from the model's internal knowledge or from a provided markdown source document; web search is not used. ## How It Works For a given topic, the tool: 1. Lists applicable subtopics using an LLM (aborts if the topic is unknown or unsupported) 2. Selects voices (single voice for monologue; male and female for duologue) via LLM 3. Generates monologue text concurrently for each subtopic 4. Deduplicates text across adjacent subtopics using a red-black odd-even approach, reducing total length by 6-40% 5. For duologues, generates dialogue text and tone instructions 6. Converts text to speech using OpenAI TTS models concurrently 7. Concatenates audio segments with `ffmpeg`, adding pauses between parts and subtopics ## Models Used - **Knowledge model** (`gpt-6-sol`): subtopic listing, voice selection, monologue text generation (from internal knowledge), duologue text generation - **Text model** (`gpt-6-sol`): monologue text generation from source documents, deduplication - **TTS model** (`gpt-4o-mini-tts-2025-12-15`): speech generation ## Usage - Requires an [OpenAI API key](https://platform.openai.com/api-keys) - Default output duration targets ~1 hour; comprehensive coverage often produces multi-hour files - Can limit sections with `--max-sections` (minimum 3) to control duration - Supports single-speaker (`--speakers 1`) and two-speaker modes - Accepts markdown source documents via `--document` - Output can include audio markers at section boundaries - Cached locally in `./work/<topic>` directory ## Configuration - Set `OPENAI_API_KEY` in `.env` file or environment - Optionally set `PODGENAI_OPENAI_MAX_WORKERS` (default: 16, max: 32) - Requires `ffmpeg` and `ffprobe` for audio concatenation ## Cost Estimation Approximate cost is $10 USD for a 10-hour podcast, with most costs from TTS generation. Shorter durations cost proportionally less. ## Output Formats - Generated MP3 files for audiobook/podcast consumption - Recommended playback speeds: 1.05x for non-technical topics, 1.0x for technical topics, 0.95x for foreign language content ## Installation Available via PyPI: `pip install podgenai` or `uv add podgenai` Can also be used as a Python library with `generate_media()` function. ## License GNU Lesser General Public License (LGPL)