Metadata-Version: 2.4
Name: cloud-ghosts
Version: 0.1.0
Summary: CLI toolkit for AWS cost optimization analysis, starting with EC2
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: boto3>=1.34.0
Requires-Dist: click>=8.1.0
Requires-Dist: rich>=13.7.0
Requires-Dist: python-dateutil>=2.8.2

# Cloud Ghosts

A Python + boto3 CLI for finding AWS cost waste, starting with EC2.
Built to grow: each service gets its own analyzer module behind a shared
`BaseAnalyzer`, so adding RDS, S3, EBS-standalone, ELB, or NAT Gateway
checks later means adding a file, not restructuring the tool.

## Install

```bash
python3 -m venv venv && source venv/bin/activate
pip install -e .
```

This installs the `cloud-ghosts` command (via `pyproject.toml`'s console_scripts entry point).

## IAM permissions needed

Read-only. Attach a policy with at minimum:

```json
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": [
      "ec2:DescribeInstances",
      "ec2:DescribeVolumes",
      "ec2:DescribeAddresses",
      "ec2:DescribeSnapshots",
      "ec2:DescribeRegions",
      "cloudwatch:GetMetricStatistics",
      "pricing:GetProducts",
      "sts:GetCallerIdentity"
    ],
    "Resource": "*"
  }]
}
```

`pricing:GetProducts` is optional -- without it the tool falls back to a small
static pricing table and logs nothing scary, it just uses the fallback silently.

## Usage

```bash
# Global options: --profile, --region, -o/--output {table,json,csv}, --output-file
cloud-ghosts --profile prod --region ap-south-1 ec2 list

# Stopped > 7 days (default). Still billed for attached EBS.
cloud-ghosts ec2 stopped --days 7

# Running instances averaging < 10% CPU over the last 14 days
cloud-ghosts ec2 underutilized --days 14 --cpu-threshold 10

# Idle EBS volumes, unattached Elastic IPs, old snapshots
cloud-ghosts ec2 volumes
cloud-ghosts ec2 eips
cloud-ghosts ec2 snapshots --days 90

# Everything at once, with an estimated total monthly savings figure
cloud-ghosts ec2 report

# Machine-readable output for piping into other tooling / dashboards
cloud-ghosts -o json --output-file report.json ec2 report
```

## What each check does

| Check | Method | Notes |
|---|---|---|
| Stopped instances | Parses `StateTransitionReason` for the stop timestamp | If AWS doesn't expose a parseable reason (e.g. stopped via some automation), the row is flagged `UNKNOWN` rather than silently dropped -- cross-check CloudTrail's `StopInstances` events for those. |
| Underutilized | CloudWatch `CPUUtilization`, 1-day granularity averaged over the window | Also reports average daily network in/out for context. Suggests one size down within the same instance family as a starting point -- not a substitute for AWS Compute Optimizer. |
| Unattached EBS volumes | `describe_volumes` filtered to `status=available` | These are billed even with nothing attached. |
| Unassociated Elastic IPs | `describe_addresses`, no `AssociationId` | AWS bills unattached EIPs hourly. |
| Old snapshots | `describe_snapshots(OwnerIds=['self'])`, age filter | Storage cost accumulates quietly; doesn't account for incremental-snapshot billing nuance. |

## Known limitations / where to extend next

- **Pricing is an estimate.** The Pricing API call assumes Linux, shared tenancy,
  no pre-installed software. Reserved Instances, Savings Plans, and Spot pricing
  aren't factored in. Treat `EstMonthlySavings` as directional, not exact --
  verify in AWS Cost Explorer before acting.
- **Right-sizing is naive.** It suggests "one size down in the same family"
  purely from the CPU average. For real workload-aware recommendations, pair
  this with **AWS Compute Optimizer** (`compute-optimizer:GetEC2InstanceRecommendations`)
  -- a good next module to add (`analyzers/compute_optimizer_analyzer.py`).
- **Single-region by default.** `--region` targets one region per run. A
  `--all-regions` flag that loops `AWSSession.all_regions()` is a natural next step.
- **Next services to add**, following the same `BaseAnalyzer` pattern:
  - RDS: idle instances (near-zero connections), oversized instances, old manual snapshots
  - EBS: gp2 volumes that should be gp3 (cheaper, faster)
  - ELB/ALB/NLB: load balancers with zero registered/healthy targets
  - NAT Gateway: low-traffic NAT Gateways that could become NAT instances or be removed
  - Lambda: over-provisioned memory vs actual usage
  - S3: lifecycle policy gaps, incomplete multipart uploads

## Project layout

```
cloud-ghosts/
  cli.py                    # click CLI, wires options -> analyzers -> output
  aws_session.py            # boto3 session/client factory, lazy credential check
  pricing.py                # live Pricing API lookup + static fallback table
  output.py                 # table (rich) / json / csv rendering
  analyzers/
    base.py                 # shared helpers (tag lookup, pagination, date math)
    ec2_analyzer.py          # all EC2 checks
```
