About this project
# podgenai
podgenai is a Python application that generates informational single-speaker or two-speaker audiobook/podcast MP3 files on a given topic using an OpenAI LLM. Content comes from the model's internal knowledge or from a provided markdown source document; web search is not used.
## How It Works
For a given topic, the tool:
1. Lists applicable subtopics using an LLM (aborts if the topic is unknown or unsupported)
2. Selects voices (single voice for monologue; male and female for duologue) via LLM
3. Generates monologue text concurrently for each subtopic
4. Deduplicates text across adjacent subtopics using a red-black odd-even approach, reducing total length by 6-40%
5. For duologues, generates dialogue text and tone instructions
6. Converts text to speech using OpenAI TTS models concurrently
7. Concatenates audio segments with `ffmpeg`, adding pauses between parts and subtopics
## Models Used
- **Knowledge model** (`gpt-6-sol`): subtopic listing, voice selection, monologue text generation (from internal knowledge), duologue text generation
- **Text model** (`gpt-6-sol`): monologue text generation from source documents, deduplication
- **TTS model** (`gpt-4o-mini-tts-2025-12-15`): speech generation
## Usage
- Requires an [OpenAI API key](https://platform.openai.com/api-keys)
- Default output duration targets ~1 hour; comprehensive coverage often produces multi-hour files
- Can limit sections with `--max-sections` (minimum 3) to control duration
- Supports single-speaker (`--speakers 1`) and two-speaker modes
- Accepts markdown source documents via `--document`
- Output can include audio markers at section boundaries
- Cached locally in `./work/<topic>` directory
## Configuration
- Set `OPENAI_API_KEY` in `.env` file or environment
- Optionally set `PODGENAI_OPENAI_MAX_WORKERS` (default: 16, max: 32)
- Requires `ffmpeg` and `ffprobe` for audio concatenation
## Cost Estimation
Approximate cost is $10 USD for a 10-hour podcast, with most costs from TTS generation. Shorter durations cost proportionally less.
## Output Formats
- Generated MP3 files for audiobook/podcast consumption
- Recommended playback speeds: 1.05x for non-technical topics, 1.0x for technical topics, 0.95x for foreign language content
## Installation
Available via PyPI: `pip install podgenai` or `uv add podgenai`
Can also be used as a Python library with `generate_media()` function.
## License
GNU Lesser General Public License (LGPL)
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.