Metadata-Version: 2.4
Name: personasurvey
Version: 0.1.0
Summary: Reusable async persona survey simulations
License-Expression: MIT
Project-URL: Homepage, https://github.com/Japulgarin/personasurvey
Project-URL: Repository, https://github.com/Japulgarin/personasurvey
Project-URL: Issues, https://github.com/Japulgarin/personasurvey/issues
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=2.0
Requires-Dist: httpx>=0.25
Requires-Dist: openai>=1.0
Requires-Dist: tqdm>=4.0
Provides-Extra: gemini
Requires-Dist: google-genai>=1.0; extra == "gemini"
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.40; extra == "anthropic"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
Dynamic: license-file

# personasurvey

Small, resumable async simulations for persona-based surveys. The package is country- and candidate-agnostic.

## Install

```bash
pip install -e .[dev]
```

For optional native providers: `pip install -e .[gemini,anthropic]`.

## Generic RunPod experiment

```python
from personasurvey import PersonaPrompt, ResponseSchema, RunPodProvider, GPUPricing, run_simulation, continue_simulation

prompt = PersonaPrompt(
    demographic_cols=["age", "region", "education", "occupation"],
    persona_detail_cols=["persona_description", "hobby"],
    scenario="How would this fictional person divide support between Candidate A and Candidate B?",
    response_schema=ResponseSchema.probabilities(["Candidate A", "Candidate B"]),
)
provider = RunPodProvider(urls=[RUNPOD_1, RUNPOD_2], model="openai/gpt-oss-20b")

prompt.validate(df)
print(prompt.preview(df.iloc[0]))

run = await run_simulation(df=df, prompt=prompt, provider=provider, checkpoint="experiment_01.jsonl", n=200_000, max_tokens=500, temperature=0, reasoning_effort="low", pricing=GPUPricing(0.53))
resumed = await continue_simulation(df=df, prompt=prompt, provider=provider, checkpoint="experiment_01.jsonl")

results_df = run.results
merged_df = run.merged
```

`continue_simulation` never sends IDs already marked `status="ok"`. Errors remain eligible for a later retry.

## Other providers and API pricing

```python
from personasurvey import APIPricing, OpenAIProvider, OpenRouterProvider, GeminiProvider, AnthropicProvider

provider = OpenAIProvider(model="gpt-4.1-mini")  # reads OPENAI_API_KEY
pricing = APIPricing(input_per_million=0.03, output_per_million=0.13)
# OpenRouterProvider(model="..."), GeminiProvider(model="..."), AnthropicProvider(model="...")
```

Credentials are supplied as arguments or environment variables; never put them in a notebook or source file.
