Production-style Linux/DevOps lab built to demonstrate hands-on configuration management, containerized application deployment, monitoring, incident response, access administration, log management, and rollback workflows.
All failures and incidents in this repository were deliberately reproduced in an isolated lab. They are not production incidents or customer data.
- 2 Ubuntu VMs on KVM/libvirt: application and monitoring nodes
- Ansible roles for baseline configuration, application deployment, developer access, Zabbix agent configuration, monitoring stack, and log rotation
- Docker + Nginx + FastAPI application path
- Zabbix + Grafana + PostgreSQL monitoring stack
- 5 documented incidents with investigation, root cause, resolution, and evidence
- Deployment health validation + rollback using immutable image tags
- GitHub Actions syntax validation for the Ansible playbooks
WSL2 Ubuntu
Ansible Control Node
|
| SSH
+----------+-----------+
| |
v v
devops-app-01 devops-monitor-01
Ubuntu 22.04 Ubuntu 22.04
192.168.122.221 192.168.122.215
| |
| +-- Zabbix Server
| +-- Zabbix Web
| +-- PostgreSQL
| +-- Grafana
|
+-- Nginx :80
+-- Docker
+-- FastAPI
+-- Zabbix Agent
+-- logrotate
Virtual machines are hosted with KVM/libvirt. The control node connects over SSH and applies the desired state with Ansible.
| Area | Technologies |
|---|---|
| Linux / virtualization | Ubuntu, KVM, libvirt, cloud-init |
| Configuration management | Ansible roles, variables, handlers |
| Application | FastAPI, Python |
| Runtime / proxy | Docker, Nginx |
| Monitoring | Zabbix Server, Zabbix Agent, Grafana |
| Data | PostgreSQL |
| Operations | SSH, systemd, logrotate, journalctl |
| Delivery validation | Ansible post-deployment health checks, GitHub Actions |
| Incident | Scenario | Detection / evidence | Recovery |
|---|---|---|---|
| INC001 | Application container outage | Zabbix HIGH alert and failed health check | Restarted container; HTTP 200 and Zabbix recovery |
| INC002 | Invalid Nginx deployment | Ansible handler failed on nginx -t |
Restored valid configuration without reloading the bad one |
| INC003 | Developer SSH access failure | Permission denied (publickey) and missing authorized_keys |
Reapplied Ansible access role and validated SSH login |
| INC004 | Excessive log growth / disk pressure | df, du, and file-size investigation |
Added Ansible-managed logrotate policy and validated rotation |
| INC005 | Broken application release | Failed deployment health check, HTTP 502, Zabbix HIGH alert | Rolled back to known-good v1; HTTP 200 and automatic alert recovery |
INC005 models a production-style release regression.
A known-good image was preserved as:
devops-demo-api:v1
A deliberately broken release was then deployed:
devops-demo-api:v2-broken
The container started, but the application listened on port 9000 while the Docker/Nginx path expected port 8000.
Observed behavior:
container running
|
+--> Uvicorn listening on :9000
|
Nginx / deployment path expects :8000
|
+--> HTTP 502
|
+--> Zabbix HIGH alert
Investigation used:
docker ps
docker inspect
docker logs
curlThe release was rolled back with the Ansible deployment playbook to devops-demo-api:v1. The health endpoint returned HTTP 200 and Zabbix automatically recorded recovery within 2 minutes.
Provides the Linux baseline across managed hosts:
- common administration packages
- operations group and user membership
- lab directories
- environment identification
Configures the application node:
- Docker and Nginx installation
- FastAPI source deployment and image build
- application container lifecycle
- Nginx reverse proxy configuration
nginx -tvalidation before reload- application health validation
Automates developer access:
- Linux account provisioning
- SSH directory and authorized key management
- access revocation
- termination of active user sessions/processes
- home-directory cleanup
Deploys:
- PostgreSQL
- Zabbix Server
- Zabbix Web
- Grafana
The database password is supplied with the ZABBIX_DB_PASSWORD environment variable and is not committed to the repository.
Configures application-host monitoring:
- Zabbix agent installation
- server and active-server settings
- runtime/log directories
- service management and restart handler
Controls application log growth:
- 50 MB rotation threshold
- five retained rotations
- compression
- delayed compression
copytruncate
Client
|
v
Nginx :80
|
v
Docker container :8000
|
v
FastAPI /health
|
v
Zabbix web scenario
|
+--> problem event
+--> recovery event
Linux host metrics are collected separately through the Zabbix agent.
linux-devops-operations-lab/
├── .github/
│ └── workflows/
│ └── ansible-validation.yml
├── ansible/
│ ├── roles/
│ │ ├── common/
│ │ ├── app_server/
│ │ ├── developer_access/
│ │ ├── monitoring_stack/
│ │ ├── zabbix_agent/
│ │ └── log_management/
│ ├── inventory.example.ini
│ ├── site.yml
│ ├── developer-access.yml
│ └── deploy-release.yml
├── app/
├── cloud-init/
├── incidents/
│ ├── INC001-application-outage/
│ ├── INC002-ansible-nginx-deployment-failure/
│ ├── INC003-developer-ssh-access-failure/
│ ├── INC004-log-growth-disk-pressure/
│ └── INC005-production-deployment-regression/
├── .env.example
├── .gitignore
└── README.md
The lab assumes:
- Linux/WSL2 control node
- Ansible
- KVM/libvirt
- two reachable Ubuntu VMs
- SSH key-based access
- Docker-capable guest systems
Copy the example inventory:
cp ansible/inventory.example.ini ansible/inventory.iniUpdate the VM IP addresses and SSH key path for your environment.
Set the monitoring database password locally:
export ZABBIX_DB_PASSWORD='replace-with-your-local-password'Do not commit the real password.
cd ansible
ansible-playbook -i inventory.ini site.yml --syntax-check
ansible-playbook -i inventory.ini site.ymlVerify connectivity:
ansible all -i inventory.ini -m pingValidate the application:
curl http://192.168.122.221/healthExpected:
{"status":"ok"}Deploy a known-good release:
cd ansible
ansible-playbook -i inventory.ini deploy-release.yml \
-e release_image=devops-demo-api:v1The playbook performs a post-deployment health check and fails the deployment workflow if the application does not return HTTP 200.
Rollback uses the same playbook with the previous known-good image tag.
Provision developer access:
ansible-playbook -i inventory.ini developer-access.ymlRevoke access:
ansible-playbook -i inventory.ini developer-access.yml \
-e developer_access_state=absentThe public-key path used by this playbook is local to the operator and is intentionally not committed.
GitHub Actions runs Ansible syntax validation on pushes and pull requests.
The CI job does not attempt to reproduce the KVM/libvirt environment or execute destructive incident scenarios. Its purpose is to catch playbook/YAML regressions before changes are merged.
The public repository intentionally excludes:
- private SSH keys
- local Ansible inventory
.envfiles- API keys
- production credentials
- hard-coded database passwords
Example files use placeholders only.
- Linux system administration
- Ansible roles, variables, handlers, and idempotent configuration
- Docker application deployment
- Nginx reverse-proxy administration
- Zabbix monitoring and alerting
- Grafana deployment
- KVM/libvirt virtualization
- Linux user and SSH access management
- application health checks
- log and disk troubleshooting
- incident investigation and root-cause analysis
- release rollback and recovery validation
This repository is a controlled technical lab built to demonstrate hands-on Linux and DevOps operations. It does not claim production ownership, customer incidents, or high-availability production experience.