AI PODCAST GENERATOR WITH VOICE CLONING
About this Gig
● Engineered an end-to-end generative media pipeline converting any topic or PDF document into a fully produced podcast episode using real cloned voices from user-uploaded personal voice samples, with a built-in AnswerBot that answers questions about the generated podcast content — achieving a Faithfulness score of 1.0 on RAGAS automated evaluation, zero hallucination on grounded queries. ● Built a 3-mode architecture — topic-only generation, PDF-grounded generation (PDF injected as LLM system prompt context for fact-rich scripts), and standalone text-to-voice — with zero-shot voice cloning via Chatterbox-TTS, global model caching and pre-computed speaker embeddings reducing per-turn audio overhead by 30–40 seconds on CPU-only hardware. ● Designed a RAG-powered AnswerBot using FAISS IndexFlatIP and SentenceTransformers (all-MiniLM-L6-v2) for sub-millisecond retrieval integrated with Groq LLaMA 3.3-70B (32–58ms inference) — users ask any question about the podcast and receive strictly context-grounded answers with RAGAS auto-scoring (Faithfulness, Relevancy, Context Precision, Recall) on every single query eliminating manual review. ● Integrated AWS S3 artifact versioning persisting scripts, embeddings, audio files, AnswerBot query logs and RAGAS scores under timestamped keys for full pipeline auditability, and resolved production-grade engineering challenges including corporate SSL proxy interception, AWS Bedrock inference profile migration, and Chatterbox watermarker binary incompatibility on Windows.
Requirements
Just the LLM api keys
