Your company knows exactly how many laptops it owns. It does not know how many of its employees have a password in circulation.
DataShield is a self-hosted service that answers that question continuously. It syncs your existing directory, checks every employee against known breach sources, and turns the result into a prioritized, auditable view of who is exposed and how badly.
v1.2.2. The production readiness checklist is tracked in docs/production-readiness.md.
- Continuous per-employee breach monitoring, not one-off lookups on a single address
- Six breach sources: Have I Been Pwned (including stealer logs), DeHashed, Intelligence X, LeakCheck and Snusbase
- Severity-based alerting with assignment, status workflow, comments and remediation tracking
Your directory stays the source of truth, so you never retype your headcount.
Six connectors, in src/lib/directory:
| Connector | Source |
|---|---|
| Microsoft Entra ID (Azure AD) | azure.ts |
| Google Workspace | google.ts |
| LDAP / Active Directory | ldap.ts |
| Okta | okta.ts |
| AWS IAM Identity Center | aws.ts |
| Inbound SCIM 2.0 | scim-auth.ts |
Connector credentials are encrypted at rest. See docs/encryption.md.
- 20 widgets, drag and drop, with saved presets and dashboards shared across
teams (see
src/lib/widgetRegistry.ts) - Breakdowns by severity, department, breach source, data type and trend over time
- GDPR exposure register with evidence attachments
- Scheduled reports in PDF, CSV and HTML
- Full audit log
- SIEM export to feed your SOC
- Outbound webhooks to your own tooling
- Data API with scoped credentials
- Better Auth with SSO (OIDC), passkeys and TOTP two-factor
- RBAC over a fixed vocabulary of 36 permissions, with role presets, step-up authentication on sensitive actions and last-admin protection
- Details in docs/auth.md
DataShield is self-hosted on purpose. Your employee data stays on your infrastructure, directory credentials are encrypted at rest, and the only outbound requests are the ones you configure. Nothing reports back to the author.
- Next.js 16 (App Router), React 19, TypeScript in strict mode
- Prisma 7 with PostgreSQL
- Better Auth (SSO, passkeys, TOTP)
- Tailwind CSS
- Vitest for unit and integration tests, Playwright for end-to-end
The quickest path needs only Docker, no Node on the host. It builds the app, starts PostgreSQL, applies migrations and seeds demo data in one command.
make env # create .env.local with generated secrets
docker compose up # app on http://localhost:3000, database seededPrerequisites: Node.js 22 (pinned via .nvmrc and engines) and Docker for the
local database, or your own PostgreSQL instance. The make targets select the
pinned Node version automatically via nvm, so they work even when your shell's
default Node differs.
make setup # install deps, create .env.local, start DB, migrate, seed
make run # start the dev servermake doctor diagnoses the setup (toolchain, env, Docker, DB, Prisma) and can
auto-fix common issues, including creating .env.local and switching to Node 22.
The underlying npm steps (npm install, npm run db:init, npm run dev) still
work if you prefer them.
Open http://localhost:3000 and sign in with the seeded admin account:
admin@datashield.local / ChangeMe123!
With docker compose up, migrations and the demo seed run on every start, so a
fresh machine is ready right after git pull.
On the host, the Prisma client is regenerated automatically on npm install
(postinstall). The database is not migrated automatically: npm run dev
only warns if migrations are pending (it never applies them on boot). After
pulling changes that add a migration, apply it explicitly before running the app:
npm run db:migrate # prisma migrate deploy, applies pending migrationsWhen you edit prisma/schema.prisma yourself, create the migration instead:
npx prisma migrate dev --name <change>npm run db:initstarts a Postgres container (compose.yml), applies all migrations and seeds demo data (breaches, employees, alerts).npm run db:up/npm run db:downstart and stop the container.npm run db:migrateapplies pending migrations to the current database.npm run seed:devreseeds the demo data;npm run seedseeds only the admin.
No Docker? Point DATABASE_URL at your own PostgreSQL, then run
npx prisma migrate deploy && npm run seed:dev.
All variables live in .env.local (copied from .env.example).
BETTER_AUTH_SECRET and DIRECTORY_ENCRYPTION_KEY must both be set to real
random values; the rest have working defaults for local development. Set
CRON_SECRET too if you want the scheduler to run.
Every push and pull request runs an automated pipeline: ESLint (zero warnings
allowed), strict type checking, Prisma schema validation and a production build,
plus CodeQL static analysis, dependency auditing, dependency review and secret
scanning. See .github/workflows. Security policy and
reporting: SECURITY.md.
Contributions are welcome.
Contribution rules are enforced automatically by Git hooks (
.githooks/) and CI (.github/workflows/compliance.yml). Non-compliant PR titles are rejected: invalid conventional commit format, AI attribution trailers, secrets, non-English text, frozen-dependency major bumps, and forbidden code patterns. The hooks activate onnpm install.
Please also read CONTRIBUTING.md and CODE_OF_CONDUCT.md.
Source-available, not open source. You may read, run, modify, fork and redistribute DataShield, including commercially, but you may not resell the software itself as a standalone product. See LICENSE for the exact terms.