Skip to content

Repository files navigation

CODEX Multitool — Multimodal Desktop Assistant

Personal project (2021) — built before ChatGPT existed, using the original OpenAI Codex.

A hands-free desktop assistant that combines computer vision, speech and early LLM integration into a single tool: control the mouse with your hand, dictate and launch commands by voice, recognize faces, detect objects in real time and query a generative model — all from one place.

🎥 Demo video: https://www.youtube.com/watch?v=SxKSmFFAkuw

Historical note: this was one of my first public personal projects, started in 2021 — over a year before ChatGPT was released. The generative module was powered by OpenAI Codex, the code-generation model that predated it. I keep the repo as-is (early code style included) because it documents where my interest in multimodal AI started; my current work on edge computer vision and NPU model deployment builds directly on what I learned here.

Modules

Module Files What it does
🖐️ Virtual mouse HandTracking.py, CVirtualMouse.py Tracks the hand through the webcam and maps gestures to cursor movement and clicks
🙂 Face recognition detectFaces.py, getFaces.py, trainFaces.py, trainFacesDLIB.py Captures faces, trains a recognizer (Haar cascades + dlib pipelines) and identifies known people in real time
🎤 Speech to text SpeechToText.py, CommandSpeech.py, subtitles.py Transcribes speech, generates live subtitles and executes voice commands
🔑 Keyword spotting getKeyword.py, trainKeyword.py, predictKeyword.py, detectKeyword.py Custom-trained wake-word/keyword detection model from .wav samples (wavsamples/)
📦 Object detection yolo4.py Real-time object detection with YOLOv4
🤖 Generative AI GenerativePretrainedTransformers.py Integration with OpenAI Codex for natural-language queries and code generation

main.py ties the modules together as a single assistant.

Stack

Python · OpenCV · dlib · YOLOv4 · OpenAI Codex API · custom audio ML pipeline

Running it

git clone https://github.com/XxSamaxX/xxxCODEXxxxV1
cd xxxCODEXxxxV1
pip install -r requirements.txt   # see note below
python main.py

⚠️ This is a 2021-era project: some dependencies (Codex API in particular) are deprecated. The computer-vision modules (virtual mouse, face recognition, YOLOv4, keyword spotting) still run standalone — each module can be launched as an individual script.

Repo layout

├── main.py                  # entry point / orchestrator
├── haarcascades/            # OpenCV cascade classifiers
├── face_recog/              # face recognition assets
├── model/                   # trained models
├── wavsamples/              # audio samples for keyword training
├── bin/ · data/             # binaries and runtime data

Built by Samuel Otero Agraso. Contributions and forks welcome.

About

Multitask tool project

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages