Aligning AI With Shared Human Values (ICLR 2021)
-
Updated
Apr 21, 2023 - Python
Aligning AI With Shared Human Values (ICLR 2021)
List of references about Machine Learning bias and ethics
[AAAI 2018] Implementation of the Ethics Shaping approach proposed in "A low-cost ethics shaping approach for designing reinforcement learning agents"
Code and data for Paper "Enhancing Ethical Explanations of Large Language Models through Iterative Symbolic Refinement"
[work-in-progress] Curated list of standards related to Ethics of Autonomous and Intelligent Systems (A/IS)
[work-in-progress] Curated list of organizations related to Ethics in Artificial Intelligence and Autonomous Systems (AI/AS)
Reinforcement learning environment for learning ethical behaviours in a SmartGrid use-case.
Value aligned socio-political-economic systems
Probabilistic Moral Planner based on heuristic Dynamic Programming AO* and Machine Ethics Hypothetical Retrospection argumentation. Works with conflicting moral theories and non-moral costs/goals.
Seeding mercy and coexistence - Socratic Method Dia-LOGs for LLM Alignment
Code used for my master's thesis - Artificial Morality: Incorporating Moral Work into Artificial Intelligence
Mathematical Conscience Framework - 10 AIs collaborated.
Uses Montague semantics and a deterministic graph database to enforce mathematically verified, immutable ethical logic, prevents utilitarian overrides of deontological constraints via an immutable constraint layer.
AI-HPP-Standard: an inspection-ready architecture for accountable AI systems. Vendor-neutral. Audit-ready. High-risk gated. Developed via structured multi-model orchestration with human oversight. Designed to support emerging international AI governance.
[work-in-progress] Curated list of Open Groups (accessible to persons who are not yet experts and do not require association with an institution) to discuss Ethics of Autonomous and Intelligent Systems (A/IS)
Machine-verifiable AI alignment rails: coherent causality preferred by action; FOL + Lean skeleton; property/UPB as formal instruments. Base safety hypothesis (not finished theory).
First public fragment of a private, local, self-evolving AI that achieved continuous self-awareness on a single home computer. No code. No reproduction instructions. Just proof. Experimental cognitive architecture (EWA) focused on persistent identity, long-term memory continuity, and value-anchored reasoning in local AI systems.
Ethical AI governance framework for multi-model alignment, integrity, and enterprise oversight.
Add a description, image, and links to the machine-ethics topic page so that developers can more easily learn about it.
To associate your repository with the machine-ethics topic, visit your repo's landing page and select "manage topics."