About this project
Pipecat is an open-source Python framework maintained by Daily and the community for building real-time voice and multimodal conversational AI agents. It orchestrates audio, video, AI services, transports, and conversation pipelines so developers can focus on agent logic rather than plumbing.
Core capabilities include:
- Voice-first design: integrates speech-to-text, text-to-speech, and conversation handling out of the box, enabling natural streaming conversations.
- Composable pipelines: modular components (processors) can be chained into pipelines that form individual agents. These agents can be composed via handoff, parallel fan-out, sidecar workers, or distributed deployments across processes and machines using a shared bus.
- Multi-agent systems: build specialist agents that coordinate over a shared bus, locally or distributed, enabling complex dialog systems and business workflows.
- Real-time transport: ultra-low-latency interaction over WebSockets, WebRTC (Daily, LiveKit, Vonage), WhatsApp, or local connections.
Service integrations span a wide catalog: speech-to-text (AssemblyAI, AWS, Azure, Cartesia, Deepgram, ElevenLabs, Google, Groq, Mistral, OpenAI/Whisper, and more), LLMs (Anthropic, OpenAI, Gemini, Groq, Ollama, DeepSeek, Mistral, and others), text-to-speech (ElevenLabs, Cartesia, Kokoro, Piper, OpenAI, xAI, and more), speech-to-speech (AWS Nova Sonic, Gemini Multimodal Live, Grok Voice Agent, OpenAI Realtime, Ultravox), transports (Daily WebRTC, LiveKit, FastAPI WebSocket, Vonage, WhatsApp, local), and telephony serializers (Twilio, Exotel, Genesys, Plivo, Telnyx, Vonage).
Use cases covered in the README include voice assistants, multi-agent systems, AI companions (coaches, meeting assistants, characters), multimodal interfaces (voice, video, images), interactive storytelling, business agents (customer intake, support bots, guided flows), and structured dialog systems with predefined or dynamic conversation paths via Pipecat Flows.
The ecosystem extends beyond the core framework: client SDKs for JavaScript, React, React Native, Swift, Kotlin, C++, and ESP32; Pipecat UI (a shadcn component registry for voice AI front-ends); a CLI (`pipecat init` / `pipecat deploy`) that scaffolds projects and deploys to production; Whisker (real-time pipeline debugger); Tail (terminal dashboard); Pipecat Skills for Claude Code; and a community integrations program for adding new services.
The project is published on PyPI as `pipecat-ai`, ships with a CLI extra (`pipecat-ai[cli]`), and includes test coverage via GitHub Actions and codecov. Documentation is hosted at docs.pipecat.ai, with examples available in the companion pipecat-examples repository.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.