Personal project (2021) — built before ChatGPT existed, using the original OpenAI Codex.
A hands-free desktop assistant that combines computer vision, speech and early LLM integration into a single tool: control the mouse with your hand, dictate and launch commands by voice, recognize faces, detect objects in real time and query a generative model — all from one place.
🎥 Demo video: https://www.youtube.com/watch?v=SxKSmFFAkuw
Historical note: this was one of my first public personal projects, started in 2021 — over a year before ChatGPT was released. The generative module was powered by OpenAI Codex, the code-generation model that predated it. I keep the repo as-is (early code style included) because it documents where my interest in multimodal AI started; my current work on edge computer vision and NPU model deployment builds directly on what I learned here.
| Module | Files | What it does |
|---|---|---|
| 🖐️ Virtual mouse | HandTracking.py, CVirtualMouse.py |
Tracks the hand through the webcam and maps gestures to cursor movement and clicks |
| 🙂 Face recognition | detectFaces.py, getFaces.py, trainFaces.py, trainFacesDLIB.py |
Captures faces, trains a recognizer (Haar cascades + dlib pipelines) and identifies known people in real time |
| 🎤 Speech to text | SpeechToText.py, CommandSpeech.py, subtitles.py |
Transcribes speech, generates live subtitles and executes voice commands |
| 🔑 Keyword spotting | getKeyword.py, trainKeyword.py, predictKeyword.py, detectKeyword.py |
Custom-trained wake-word/keyword detection model from .wav samples (wavsamples/) |
| 📦 Object detection | yolo4.py |
Real-time object detection with YOLOv4 |
| 🤖 Generative AI | GenerativePretrainedTransformers.py |
Integration with OpenAI Codex for natural-language queries and code generation |
main.py ties the modules together as a single assistant.
Python · OpenCV · dlib · YOLOv4 · OpenAI Codex API · custom audio ML pipeline
git clone https://github.com/XxSamaxX/xxxCODEXxxxV1
cd xxxCODEXxxxV1
pip install -r requirements.txt # see note below
python main.py
⚠️ This is a 2021-era project: some dependencies (Codex API in particular) are deprecated. The computer-vision modules (virtual mouse, face recognition, YOLOv4, keyword spotting) still run standalone — each module can be launched as an individual script.
├── main.py # entry point / orchestrator
├── haarcascades/ # OpenCV cascade classifiers
├── face_recog/ # face recognition assets
├── model/ # trained models
├── wavsamples/ # audio samples for keyword training
├── bin/ · data/ # binaries and runtime data
Built by Samuel Otero Agraso. Contributions and forks welcome.