Metadata-Version: 2.5
Name: async-site-crawler
Version: 0.1.0
Summary: High-performance async concurrent web crawler that recursively crawls websites and exports structured data to CSV
Project-URL: Homepage, https://github.com/Utkarsh736/webcrawler
Project-URL: Repository, https://github.com/Utkarsh736/webcrawler
Project-URL: Issues, https://github.com/Utkarsh736/webcrawler/issues
Author-email: Utkarsh736 <utkarshtomar736@gmail.com>
License: MIT
License-File: LICENSE
Keywords: aiohttp,async,asyncio,crawler,scraping,web-crawler
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.12
Requires-Dist: aiohttp>=3.9
Requires-Dist: beautifulsoup4>=4.12
Provides-Extra: dev
Requires-Dist: requests; extra == 'dev'
Description-Content-Type: text/markdown

# async-site-crawler

A high-performance, concurrent web crawler built in Python that recursively crawls websites and exports structured data to CSV reports.

## Features

- **Async/Concurrent Crawling**: Uses `asyncio` and `aiohttp` for fast, non-blocking HTTP requests
- **Configurable Limits**: Control max concurrent requests and total pages to crawl
- **Smart URL Handling**: Normalizes URLs to avoid duplicate crawls
- **Same-Domain Filtering**: Stays within the target website domain
- **HTML Parsing**: Extracts h1 tags, paragraphs, links, and images using BeautifulSoup
- **CSV Export**: Generates structured reports for easy analysis
- **Graceful Stopping**: Cancels in-flight tasks when limits are reached
- **Error Handling**: Handles timeouts, non-HTML content, and network failures

## Installation

```bash
pip install async-site-crawler
```
Or with uv:
```Bash
uv add async-site-crawler
```

## Usage
### Command Line
```Bash
async-site-crawler https://example.com
```
With custom settings:
```Bash
async-site-crawler https://example.com 10 50
```


| Argument          | Description                      | Default |
| ----------------- | -------------------------------- | ------- |
| `URL`             | Starting URL to crawl (required) | \-      |
| `max_concurrency` | Maximum concurrent HTTP requests | 5       |
| `max_pages`       | Maximum number of pages to crawl | 100     |


## As a Python Library
```Python
import asyncio
from async_site_crawler import crawl_site_async, write_csv_report

async def main():
    page_data = await crawl_site_async(
        base_url="https://example.com",
        max_concurrency=5,
        max_pages=20
    )
    write_csv_report(page_data, filename="report.csv")

asyncio.run(main())
```

## Output
The crawler generates a `report.csv` file with the following columns:

- **page_url**: The URL of the crawled page
- **h1**: The main heading (h1 tag) content
- **first_paragraph**: Text from the first paragraph (prioritizes `<main>` tag)
- **outgoing_link_urls**: All links found on the page (semicolon-separated)
- **image_urls**: All image URLs found on the page (semicolon-separated)

## License
This project is open source and available under the **MIT License**.

## Author
[Utkarsh736](https://github.com/Utkarsh736)
