A full-stack application that transforms text-based documentary/storytelling scripts into professional video segments using AI-powered image search and audio synchronization.
It takes an audio recording and text input, segments the text using AI, aligns the text with audio via Whisper, automatically searches Google Images, and produces perfectly lip-synced MP4 video segments for each part.
- Features
- Architecture
- Demo Flow
- Quick Start (Docker)
- Manual Setup (Development)
- API Keys
- Project Structure
- API Reference
- Technical Details
- Troubleshooting
- License
| Feature | Description |
|---|---|
| 🤖 AI Text Segmentation | Segments documentary text into meaningful parts using Gemini API |
| 🎙️ Whisper Audio Analysis | Extracts word-level timestamps using OpenAI Whisper (GPU-accelerated) |
| 🔍 Automatic Image Search | Searches for contextually relevant images via Google Images for each segment |
| 🎬 Video Generation | Creates perfectly synchronized MP4 segments using FFmpeg |
| 🎯 Frame-Aligned Sync | 30 FPS frame alignment guarantees zero cumulative audio drift |
| 📱 Modern Web UI | Responsive, dark-mode interface built with React + Vite |
| 🐳 Docker Ready | Single-command deployment with NVIDIA GPU support |
| 💾 Session Management | Restores session state even after page reloads |
┌─────────────────────────────────────────────────────────┐
│ Docker Container │
│ │
│ ┌──────────┐ ┌────────────────────────────────┐ │
│ │ Nginx │────▶│ React Frontend (Static) │ │
│ │ :80 │ │ - UploadView │ │
│ │ │ │ - ReviewView │ │
│ │ │ │ - StudioView │ │
│ │ │ └────────────────────────────────┘ │
│ │ │ │
│ │ │ ┌────────────────────────────────┐ │
│ │ /api/* ─┼────▶│ FastAPI Backend (Uvicorn) │ │
│ └──────────┘ │ :8005 │ │
│ │ │ │
│ │ ┌───────────┐ ┌─────────────┐ │ │
│ │ │ Gemini AI │ │ Whisper │ │ │
│ │ │ (Text) │ │ (Audio→Text) │ │ │
│ │ └───────────┘ └──────┬──────┘ │ │
│ │ │ │ │
│ │ ┌───────────┐ ┌──────▼──────┐ │ │
│ │ │ Serper │ │ NVIDIA GPU │ │ │
│ │ │ (Images) │ │ (CUDA) │ │ │
│ │ └───────────┘ └─────────────┘ │ │
│ │ │ │
│ │ ┌─────────────────────────┐ │ │
│ │ │ FFmpeg (Video Output) │ │ │
│ │ └─────────────────────────┘ │ │
│ └────────────────────────────────┘ │
└─────────────────────────────────────────────────────────┘
Upload your audio file (MP3/WAV) and documentary script. The AI automatically segments the text and aligns it with the audio using Whisper.
Check the AI's audio-text alignment. Edit and confirm any problematic matches.
Select one of the images retrieved from Google Images for each segment. Once selected, the video is generated and downloaded automatically.
Import the downloaded MP4 segments in sequence into your video editor (CapCut, DaVinci Resolve, etc.). The audio synchronization will be flawless.
| Component | Minimum | Recommended |
|---|---|---|
| Docker Engine | 24.0+ | Latest Version |
| RAM | 4 GB | 8 GB+ |
| Disk | 10 GB | 20 GB+ |
| GPU | - | NVIDIA (CUDA 12.x) |
# 1. Install NVIDIA Container Toolkit (first time only)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# 2. Clone the project
git clone https://github.com/YOUR_USERNAME/google_images_to_video.git
cd google_images_to_video
# 3. Setup API keys
cp backend/.env.example backend/.env
nano backend/.env # Enter your API keys
# 4. Build and run
docker compose up -d --build
# 5. Open in browser
# http://localhost# Follow the same steps, but use this command for step 4:
docker compose -f docker-compose.cpu.yml up -d --build
⚠️ Whisper runs ~3-5x slower in CPU mode, but all functions work smoothly.
- Python 3.10+
- Node.js 18+
- FFmpeg
- (Optional) NVIDIA GPU + CUDA
# Create virtual environment
python3 -m venv venv
source venv/bin/activate
# Install dependencies
cd backend
pip install -r requirements.txt
# PyTorch CUDA (if GPU is available)
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126
# Setup API keys
cp .env.example .env
nano .env
# Start the server
uvicorn app.main:app --host 0.0.0.0 --port 8005cd frontend
npm install
npm run devOpen http://localhost:3000 in your browser.
This application requires 2 external API services:
| Service | Purpose | Free Tier | Where to Get |
|---|---|---|---|
| Gemini API | Text segmentation & translation | ✅ 1500 req/day | Google AI Studio |
| Serper API | Google Images search | ✅ 2500 req/month | serper.dev |
Add your API keys to the backend/.env file:
GEMINI_API_KEY=your_gemini_api_key_here
SERPER_API_KEY=your_serper_api_key_heregoogle_images_to_video/
├── backend/
│ ├── app/
│ │ ├── main.py # FastAPI endpoints & session management
│ │ └── services/
│ │ ├── text_to_json.py # Text segmentation via Gemini AI
│ │ ├── audio_processing.py # Audio-text alignment via Whisper
│ │ ├── image_search.py # Image search via Serper API
│ │ └── video_generation.py # Video generation via FFmpeg
│ ├── system_prompt.txt # Gemini AI system instructions
│ ├── requirements.txt # Python dependencies
│ ├── .env.example # Example environment variables
│ └── .env # API keys (not tracked by git)
│
├── frontend/
│ ├── src/
│ │ ├── App.jsx # Main application component
│ │ ├── components/
│ │ │ ├── UploadView.jsx # Audio & text upload view
│ │ │ ├── ReviewView.jsx # Timestamp review view
│ │ │ └── StudioView.jsx # Image selection & video generation view
│ │ ├── index.css # Global styles (dark theme)
│ │ └── main.jsx # React entry point
│ ├── package.json
│ └── vite.config.js # Vite config & API proxy settings
│
├── Dockerfile # Multi-stage build (React + CUDA Python)
├── docker-compose.yml # GPU supported deployment
├── docker-compose.cpu.yml # CPU-only deployment
├── nginx.conf # Reverse proxy configuration
├── docker-entrypoint.sh # Container startup script
├── .gitignore
├── .dockerignore
├── LICENSE
└── README.md
| Method | Endpoint | Description |
|---|---|---|
POST |
/api/analyze |
Upload audio + text, start AI analysis |
GET |
/api/session/{id} |
Get session details |
POST |
/api/session/{id}/confirm |
Confirm timestamps |
GET |
/api/session/{id}/images |
Get images for the current segment |
POST |
/api/session/{id}/generate_video |
Generate video with the selected image |
POST |
/api/session/{id}/next |
Skip to next segment |
POST |
/api/session/{id}/prev |
Go back to previous segment |
POST |
/api/session/{id}/next_keyword |
Search with a different keyword |
GET |
/api/session/healthcheck |
Health check endpoint |
Detailed API documentation: http://localhost:8005/docs (Swagger UI)
This application guarantees zero cumulative drift between video segments:
- Frame-Aligned Timestamps: All timestamps are aligned to exact multiples of
1/30 seconds. - WAV Extraction: Audio is extracted to PCM/WAV format first (sample-accurate cuts).
- Matched Duration: Video and audio tracks are produced with exactly the same duration.
- Sequential Boundary Sync: Boundary points of consecutive segments are perfectly aligned.
By default, the small model is used. For higher accuracy, medium or large-v3 can be used (via the model_name parameter in audio_processing.py).
| Model | VRAM | Speed | Accuracy |
|---|---|---|---|
| tiny | ~1 GB | ⚡⚡⚡⚡ | ★★☆☆ |
| small | ~2 GB | ⚡⚡⚡ | ★★★☆ |
| medium | ~5 GB | ⚡⚡ | ★★★★ |
| large-v3 | ~10 GB | ⚡ | ★★★★★ |
Sources likely to contain watermarks (Getty Images, Shutterstock, iStock, etc.) are automatically filtered out from the image search results.
Docker build takes too long
The initial build may take ~10-15 minutes (PyTorch + CUDA download). Subsequent builds are much faster thanks to the Docker cache. You can run docker compose build --no-cache to force a complete rebuild.
GPU is not detected
# Check NVIDIA driver
nvidia-smi
# Is NVIDIA Container Toolkit installed?
dpkg -l | grep nvidia-container-toolkit
# Is the Docker runtime correct?
docker info | grep -i runtime"The source site for this image does not allow downloading" error
Some websites block bot access. Select a different image or use the 🔑 button to try a different keyword.
Whisper is running very slowly
Whisper is slow in CPU mode. If you have an NVIDIA GPU, make sure you are using docker-compose.yml (the GPU version). Using a GPU provides a ~5-10x speed boost.
Page refreshes return to the upload screen
In-memory sessions are cleared if the backend server restarts. Make sure the server is running stably: docker compose ps
This project is licensed under the MIT License.
Built with ❤️ using Python, React, Whisper, Gemini & FFmpeg