Metadata-Version: 2.4
Name: adr-enforcer
Version: 0.1.0
Summary: Deterministic Tree-sitter AST enforcement of Architecture Decision Records, with RAG-grounded LLM auto-fix
Author-email: Sai Phani Krishna Arumalla <saiphanikrishna05@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/Saiphanikrishna05/ArchGuard
Project-URL: Repository, https://github.com/Saiphanikrishna05/ArchGuard
Project-URL: Issues, https://github.com/Saiphanikrishna05/ArchGuard/issues
Keywords: architecture,adr,linter,tree-sitter,static-analysis,langgraph,rag,code-review
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Software Development :: Libraries
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: tree-sitter>=0.20.0
Requires-Dist: tree-sitter-javascript
Requires-Dist: tree-sitter-typescript
Requires-Dist: langgraph
Requires-Dist: langchain-core
Requires-Dist: langchain-groq
Requires-Dist: pydantic
Requires-Dist: chromadb
Requires-Dist: typer
Requires-Dist: rich
Requires-Dist: python-dotenv
Requires-Dist: mcp
Dynamic: license-file

# ARCH GUARD: Agentic Architectural Governance

## The Problem
As software systems evolve, the divergence between intended design and the actual codebase—known as architectural drift—silently accumulates massive technical debt. Traditional AI coding assistants often exacerbate this issue; they generate code that passes unit tests but violates complex, system-wide architectural constraints because they lack structural awareness. Furthermore, standard RAG pipelines rely on naive text chunking, which destroys the semantic boundaries of code.

## The Solution
This project is an enterprise-grade, multi-agent orchestration system that enforces Architecture Decision Records (ADRs) directly within the pull request lifecycle. By moving architectural governance out of static wikis and into the CI/CD workflow, this tool ensures that code and architecture stay in sync by construction.

## Core Engineering Architecture

- **ADR-Driven Rule Extraction**: ADRs live as markdown in `./adrs`. A Groq-backed extractor (`--extract`) parses each one into a structured `ExtractedRule` (with a `pattern_type`) and caches it to `data/extracted_rules.json`. Adding a new governed rule is open/closed — write the ADR, register one AST matcher for its `pattern_type`, and it is enforced. Ships enforcing three rules: **ADR-003** (no default exports), **ADR-005** (no `console.*`), **ADR-007** (no `any` type).

- **Deterministic AST Detection**: Bypasses unreliable string-matching by using Tree-sitter to parse each file into an Abstract Syntax Tree and matching exact node structures. Detection runs before any LLM is invoked, so it is precise, language-aware, and repeatable (zero false positives on the golden set).

- **Retrieval-Augmented Refactoring (RAG)**: The codebase is indexed into a ChromaDB vector store as **whole AST nodes** (complete functions/classes/arrow bindings) rather than naive text chunks, preserving semantic boundaries. When generating a fix, the agent retrieves the most relevant *already-compliant* blocks from the same repo and uses them as few-shot context, so rewrites match the project's real house style.

- **Planner-Executor Multi-Agent Pattern**: Built with LangGraph, this system decouples reasoning from action — a Planner generates an ordered fix checklist, an Executor applies one step at a time, and a Re-planner reviews and terminates (with a retry circuit-breaker), reducing hallucinations and context bloat. The `fix` path turns this into **applied** changes: it rewrites violating files, verifies the rewrite actually removes the violations via a re-scan, and writes them to disk (`--apply`) or shows a dry-run diff.

- **Quantitative Evaluation Harness**: Benchmarked against a custom "Golden Dataset" of seeded violations. Detection is a deterministic correctness gate (Precision/Recall/F1); the LLM is only exercised for refactoring, measured with Pass@k and verified by re-parsing the output with Tree-sitter — not "vibes-based" testing.

- **Production Deployment**: Operates as a GitHub Action composite workflow (scans the PR, comments findings, fails the check) and runs locally via a Model Context Protocol (MCP) Python server for native integration with AI IDEs like Cursor and Claude Desktop.

## Evaluation Metrics
To validate production viability, the system is put through an automated `harness.py` testing suite against a curated "Golden Dataset" of valid code and seeded violations (**16 cases across all three rules — ADR-003, ADR-005, ADR-007 — 10 violations, 6 valid**). Each case is scoped to the single ADR it exercises. Detection is deterministic (Tree-sitter AST), so it flags real violations without hallucinated false positives; the LLM refactor output is then verified both by re-parsing it with Tree-sitter and by re-running the detector to confirm the violation is actually gone (pass@k).

**Latest measured results** (reproduce with `python run_evaluation.py`, written to `evaluation_report.json`):

| Metric | Score |
|---|---|
| Precision | 1.00 |
| Recall | 1.00 |
| F1-Score | 1.00 |
| Pass@1 | 1.00 |
| Pass@3 | 1.00 |

Because detection is deterministic AST matching (not a probabilistic model), perfect Precision/Recall on this curated set is *expected* — the harness functions as a **correctness gate** that fails CI if the detector ever regresses, not as an ML performance benchmark. The interesting signal is Pass@k, which measures whether the LLM's refactor is syntactically valid and actually removes the violation.

> **LLM provider:** the refactoring agents and the ADR rule extractor run on **Groq** (`openai/gpt-oss-120b`). Set `GROQ_API_KEY` in your environment or a `.env` file. Note that **scanning/detection is fully deterministic and needs no API key** — only the AI-generated refactoring step calls the LLM.

## Usage Instructions

### 1. Command Line Execution (CLI)
You can run the tool locally.
```bash
# Install the package
pip install .

# Scan a local repository or file (exits non-zero if violations are found)
adr-enforcer scan ./src/components

# Derive the rules live from your ADR markdown in ./adrs, then scan
adr-enforcer scan ./src/components --extract

# Auto-fix: preview the compliant rewrite as a diff (dry-run) ...
adr-enforcer fix ./src/components

# ... or apply it to disk (verified to remove the violations first)
adr-enforcer fix ./src/components --apply

# Use as a PR/CI gate
adr-enforcer review-pr --pr-number 42 --path .
```

### 2. GitHub Composite Action
Enforce ADRs continuously in CI/CD by utilizing the wrapped GitHub Action on Pull Requests. The action scans the checked-out code, posts a findings comment on the PR, and fails the check if any violation is found:

```yaml
steps:
  - uses: actions/checkout@v4
  - name: Run ADR Enforcer
    uses: your-username/adr-enforcer@main
    with:
      github_token: ${{ secrets.GITHUB_TOKEN }}
      groq_api_key: ${{ secrets.GROQ_API_KEY }}   # optional; scanning works without it
```

### 3. Model Context Protocol (MCP) Server
Integrate directly into your AI-native IDE (Cursor / Windsurf / Claude Desktop). Put this into your `mcp.json` settings:
```json
{
  "mcpServers": {
    "adr-enforcer": {
      "command": "python",
      "args": ["-m", "adr_enforcer.mcp_server"]
    }
  }
}
```
This exposes two tools to the IDE's local AI agent: `scan_architectural_compliance` (detect violations) and `apply_architectural_fixes` (generate/apply compliant rewrites) against your workspace on demand.

## Defining Your Own ADRs
Drop a markdown file in `./adrs` describing the decision, then map it to a matcher:

1. Write the ADR (see `adrs/ADR-003-no-default-exports.md` for the format).
2. Add a `PatternType` value in `adr_enforcer/schemas/rule_schema.py`.
3. Register a Tree-sitter matcher for it in `adr_enforcer/detector.py` (`MATCHERS`).
4. Run `adr-enforcer scan <path> --extract` — the LLM extracts the structured rule and the detector enforces it.

Rules whose `pattern_type` has no matcher yet are reported as "recognised but not-yet-implemented" rather than silently ignored.
