Fyrefly
v0.6.0

Fyrefly

A quant-based Python library for flexible math, financial formulas, equation solving, SQL-powered data analysis, AI-driven insights, and visualization — every function takes any number of arguments.

Installation

Fyrefly is a single package — every function in this reference comes from one install.

$pip install fyrefly

Quickstart

Most real usage touches several parts of the library in one script. Here's a realistic end-to-end example — load a file, clean it, query it, chart it, and ask a question about it in plain English.

from fyrefly import load, clean, sql, viz, ask

# Load a dataset — CSV, Excel, or a Google Sheet all work
c = load("sales.csv")

# Clean it up: dedupe, strip whitespace, standardize column names
c = clean(c)

# Query it with real SQL, by the variable name you gave it
top_regions = sql("select region, sum(amount) as total from c group by region order by total desc")

# Visualize the result — one line, no separate charting library
viz.bar("region", "amount")

# Or just ask in plain English (needs your own API key)
ask(c, "which region grew the fastest last quarter?", api_key="AIza...")

Everything below is organized by what it does. Math, sequences, and geometry need nothing but Python. Data operations need a file to load. The AI functions are the only ones that need an API key of your own — everything else runs completely offline.

Math

The building blocks. Each function takes any number of arguments and applies the operation left to right — add(1, 2, 3, 4) sums all four; div(100, 5, 2) divides 100 by 5, then that result by 2.

add(*nos)

Adds any count of numbers.

add(5, 4, 6, 10)
# 25

add(1, 2, 3, 4, 5)
# 15

mul(*nos)

Multiplies any count of numbers.

mul(5, 4, 6, 10)
# 1200

div(*nos)

Divides numbers in sequence: first ÷ second ÷ third ...

div(100, 5, 2)
# 10.0

mod(*nos)

Finds the remainder in sequence: first % second % third ...

mod(17, 5)
# 2

diff(*nos)

Subtracts numbers in sequence: first − second − third ...

diff(10, 3, 2)
# 5

Quant & Finance

Interest and profit/loss formulas for everyday financial math. profit() and loss() are deliberately strict: call the wrong one — say profit() on numbers that are actually a loss — and you get a clear error telling you which function to use instead, rather than a silently wrong (negative) answer.

si(p, n, r)

Simple Interest: (p × n × r) / 100.

si(1000, 2, 5)
# 100.0

ci(p, n, r, freq="annual")

Compound interest earned (principal excluded). freq is "annual" or "semi-annual".

ci(1000, 2, 10)
# 210.0

ci(1000, 2, 10, "semi-annual")
# 215.5

si_ci_diff(p, r, n)

Difference between Compound and Simple Interest. Supports n = 2 or 3 years only.

si_ci_diff(1000, 10, 2)
# 10.0

si_ci_diff(1000, 10, 3)
# 31.0

profit(sp, cp, pct=None)

Profit = sp − cp. Pass "%" as the third argument for (value, percent). Raises a clear error if sp < cp — that's a loss, not a profit.

profit(120, 100)
# 20

profit(120, 100, "%")
# (20, 20.0)

profit(80, 100)
# raises: this is a LOSS. Use loss(80, 100) instead.

loss(sp, cp, pct=None)

Loss = cp − sp. Pass "%" as the third argument for (value, percent). Raises a clear error if sp > cp — that's a profit, not a loss.

loss(80, 100)
# 20

loss(80, 100, "%")
# (20, 20.0)

Sequences

Each progression type works the same way: call it directly for the full list of terms, or use .tn() and .sn() when you only need one term or a running total — no need to generate the whole sequence first.

ap(a, d, n) · ap.tn(a, d, n) · ap.sn(a, d, n)

Arithmetic Progression: full sequence, the nth term, or the sum of the first n terms.

ap(2, 3, 4)
# [2, 5, 8, 11]

ap.tn(2, 3, 4)
# 11

ap.sn(2, 3, 4)
# 26.0

gp(a, r, n) · gp.tn(a, r, n) · gp.sn(a, r, n)

Geometric Progression: full sequence, the nth term, or the sum of the first n terms.

gp(2, 3, 4)
# [2, 6, 18, 54]

gp.tn(2, 3, 4)
# 54

gp.sn(2, 3, 4)
# 80.0

hp(a, d, n) · hp.tn(a, d, n) · hp.sn(a, d, n)

Harmonic Progression — reciprocals of an AP with first term 1/a. No closed-form sum exists, so .sn() totals the actual terms.

hp(2, 3, 4)
# [2.0, 0.2857..., 0.1538..., 0.1052...]

hp.tn(2, 3, 4)
# 0.10526315789473684

Averages & Statistics

avg() takes raw numbers directly; avgc() is for when you already know a value's frequency rather than listing it out repeatedly — e.g. avgc(5, 3) for "the value 5, occurring 3 times", instead of avg(5, 5, 5).

avg(*nos)

Plain average of individual numbers.

avg(10, 12, 13)
# 11.666666666666666

avgc(*value_count_pairs)

Weighted average from (value, count) pairs — e.g. "this value occurred this many times."

avgc(5, 3)
# 5.0   (three 5's)

avgc(5, 3, 10, 2)
# 7.0   (three 5's and two 10's)

mode(*nos)

Most frequently occurring value(s). Returns a list — may contain more than one value if there's a tie.

mode(1, 2, 2, 3)
# [2]

mode(1, 1, 2, 2)
# [1, 2]

median(*nos)

Median of any count of numbers.

median(1, 3, 2)
# 2

median(1, 2, 3, 4)
# 2.5

Geometry

Coordinate geometry formulas you'd otherwise look up every time. Points are plain (x, y) tuples throughout, so these compose easily with your own data.

hyp(a, b)

Hypotenuse of a right triangle given the other two sides.

hyp(3, 4)
# 5.0

slope(point1, point2)

Slope between two (x, y) tuples: (y2 − y1) / (x2 − x1). Raises a clear error for a vertical line.

slope((1, 2), (3, 4))
# 1.0

centroid(point1, point2, point3)

Centroid of a triangle given three (x, y) tuples.

centroid((1, 2), (3, 2), (4, 5))
# (2.6666666666666665, 3.0)

dist(point1, point2)

Distance between two (x, y) tuples: √((x2−x1)² + (y2−y1)²).

dist((6, 2), (10, 5))
# 5.0

Mensuration — 2D Shapes

Area and perimeter for every common 2D shape. Each shape is its own object with .area() and .perimeter() (or close variants) as methods.

circle.area(r) · circle.perimeter(r)

Area and circumference of a circle.

circle.area(7)
# 153.93804002589985

circle.perimeter(7)
# 43.982297150257104

semicircle.area(r) · semicircle.perimeter(r)

Area and perimeter (curved edge + diameter) of a semicircle.

semicircle.area(7)
# 76.96902001294993

semicircle.perimeter(7)
# 35.99114857512855

square.area(s) · square.perimeter(s)

Area and perimeter of a square.

square.area(5)
# 25

square.perimeter(5)
# 20

rectangle.area(l, w) · rectangle.perimeter(l, w)

Area and perimeter of a rectangle.

rectangle.area(4, 6)
# 24

rectangle.perimeter(4, 6)
# 20

triangle.area(base, height) · triangle.area_sides(a, b, c) · triangle.perimeter(a, b, c)

Area from base/height, area from all three sides via Heron's formula, and perimeter. Raises a clear error if the three sides can't form a real triangle.

triangle.area(10, 5)
# 25.0

triangle.area_sides(3, 4, 5)
# 6.0

triangle.perimeter(3, 4, 5)
# 12

equilateral_triangle.area(s) · equilateral_triangle.perimeter(s)

Area and perimeter of an equilateral triangle from its side length.

equilateral_triangle.area(4)
# 6.928203230275509

equilateral_triangle.perimeter(4)
# 12

parallelogram.area(base, height) · parallelogram.perimeter(a, b)

Area and perimeter of a parallelogram.

parallelogram.area(10, 5)
# 50

parallelogram.perimeter(10, 5)
# 30

rhombus.area(d1, d2) · rhombus.perimeter(s)

Area from the two diagonals, and perimeter from the side length.

rhombus.area(6, 8)
# 24.0

rhombus.perimeter(5)
# 20

kite.area(d1, d2) · kite.perimeter(a, b)

Area from the two diagonals, and perimeter from the two distinct side lengths.

kite.area(6, 8)
# 24.0

kite.perimeter(5, 7)
# 24

trapezoid.area(a, b, height) · trapezoid.perimeter(a, b, c, d)

Area from the two parallel sides and height, and perimeter from all four sides.

trapezoid.area(8, 5, 4)
# 26.0

trapezoid.perimeter(8, 5, 3, 4)
# 20

ellipse.area(a, b) · ellipse.perimeter(a, b)

Area is exact. Perimeter uses Ramanujan's approximation — extremely accurate, since no exact closed-form formula exists for an ellipse's perimeter.

ellipse.area(5, 3)
# 47.12388980384689

ellipse.perimeter(5, 3)
# ~25.5 (approximate)

polygon.area(n_sides, length) · polygon.perimeter(n_sides, length)

Area and perimeter of any regular polygon, given its number of sides and side length.

polygon.area(6, 4)
# area of a regular hexagon

polygon.perimeter(6, 4)
# 24

sector.area(r, angle) · sector.arc_length(r, angle)

Area and arc length of a circular sector. Angle is in degrees.

sector.area(7, 90)
# 38.48451000647496

sector.arc_length(7, 90)
# 10.995574287564276

annulus.area(outer_r, inner_r)

Area of a ring shape — the area between two concentric circles.

annulus.area(10, 6)
# 201.06192982974676

Mensuration — 3D Shapes

Volume and surface area for every common 3D shape. Surface area is split into .tsa() (Total Surface Area, every face combined), .csa() (Curved Surface Area, curved part only — for shapes with a genuine curve), and .lsa() (Lateral Surface Area, side faces only — for flat-sided shapes). Sphere and Torus use plain .surface_area() instead, since there's nothing to disambiguate on those two.

cube.volume(s) · cube.lsa(s) · cube.tsa(s)

Volume, lateral surface area (4 side faces), and total surface area (all 6 faces) of a cube.

cube.volume(3)
# 27

cube.lsa(3)
# 36

cube.tsa(3)
# 54

cuboid.volume(l, w, h) · cuboid.lsa(l, w, h) · cuboid.tsa(l, w, h)

Volume, lateral surface area, and total surface area of a cuboid.

cuboid.volume(2, 3, 4)
# 24

cuboid.lsa(2, 3, 4)
# 40

cuboid.tsa(2, 3, 4)
# 52

sphere.volume(r) · sphere.surface_area(r)

Volume and surface area of a sphere. No TSA/CSA split — the whole surface is uniform.

sphere.volume(3)
# 113.09733552923254

sphere.surface_area(3)
# 113.09733552923255

hemisphere.volume(r) · hemisphere.csa(r) · hemisphere.tsa(r)

Volume, curved surface area (dome only), and total surface area (dome + flat base) of a hemisphere.

hemisphere.volume(3)
# 56.54866776461627

hemisphere.csa(3)
# 56.548667764616276

hemisphere.tsa(3)
# 84.82300164692441

cylinder.volume(r, h) · cylinder.csa(r, h) · cylinder.tsa(r, h)

Volume, curved surface area (side only), and total surface area (side + both circular ends) of a cylinder.

cylinder.volume(3, 7)
# 197.92033717615698

cylinder.csa(3, 7)
# 131.94689145077132

cylinder.tsa(3, 7)
# 188.49555921538757

cone.volume(r, h) · cone.csa(r, h) · cone.tsa(r, h)

Volume, curved surface area, and total surface area of a cone. Slant height is computed internally from r and h.

cone.volume(3, 4)
# 37.69911184307752

cone.csa(3, 4)
# 47.12388980384689 (slant height = 5)

cone.tsa(3, 4)
# 75.39822368615503

frustum.volume(r1, r2, h) · frustum.csa(r1, r2, h) · frustum.tsa(r1, r2, h)

Volume, curved surface area, and total surface area of a frustum (a cone with the tip cut off). Slant height is computed internally.

frustum.volume(3, 5, 6)
# 307.8760800517997

frustum.csa(3, 5, 6)
# 158.95341225273762

frustum.tsa(3, 5, 6)
# 265.76756247479057

pyramid.volume(s, h) · pyramid.lsa(s, h) · pyramid.tsa(s, h)

Square-based pyramid. h is the vertical height — slant height for the triangular faces is computed internally.

pyramid.volume(6, 8)
# 96.0

pyramid.lsa(6, 8)
# 102.52804494381036

pyramid.tsa(6, 8)
# 138.52804494381036

prism.volume(base_area, h) · prism.lsa(base_perimeter, h) · prism.tsa(base_area, base_perimeter, h)

General prism of any base shape — described by its base area and base perimeter, rather than a specific shape.

prism.volume(20, 10)
# 200

prism.lsa(18, 10)
# 180

prism.tsa(20, 18, 10)
# 220

torus.volume(R, r) · torus.surface_area(R, r)

Volume and surface area of a torus (donut shape). R is the distance from the tube's center to the torus's center; r is the tube's own radius. No TSA/CSA split — the whole surface is uniform.

torus.volume(5, 2)
# 394.78417604357435

torus.surface_area(5, 2)
# 394.78417604357435

Equations

eqn.l() solves a system of simultaneous equations — give it one tuple per equation (its coefficients, then the constant) and it works out every unknown. eqn.q() finds the roots of a polynomial of any degree, including complex roots when there's no real solution.

eqn.l(*equations)

Solves n linear equations in n unknowns. Each equation is (coefficients..., constant) — n+1 numbers for n unknowns.

eqn.l((2, 3, 5), (7, 6, 10))
# (x, y) solving 2x+3y=5, 7x+6y=10

eqn.l((1,1,1,6), (2,-1,1,3), (1,2,-1,2))
# (a, b, c) — scales to any n

eqn.q(*coeffs) · eqn.q.sum(*coeffs) · eqn.q.mul(*coeffs)

Roots of a polynomial of any degree, plus the sum and product of all roots (Vieta's formulas — valid at any degree).

eqn.q(1, 3, 4)
# roots of x^2 + 3x + 4 = 0

eqn.q(3, 4, 6, 9, 10)
# roots of the quartic

eqn.q.sum(5, 2, 3)
# -0.4

eqn.q.mul(5, 2, 3)
# 0.6

Logs, Exponents & Roots

Everyday algebra helpers — logarithms with any base, exponents in either direction, and roots of any degree, including correct handling of negative numbers for odd roots.

log(x, base=e)

Logarithm of x. Natural log by default; pass a second argument for a different base.

log(math.e)
# 1.0

log(8, 2)
# 3.0

exp(a, b)

a raised to the power b. Works for positive and negative exponents.

exp(2, 5)
# 32

exp(2, -1)
# 0.5

sqrt(no) · curt(no) · nroot(no, n)

Square root, cube root, and the generic nth root. Odd roots of negative numbers work correctly.

sqrt(16)
# 4.0

curt(27)
# 3.0

curt(-27)
# -3.0

nroot(16, 2)
# 4.0

Trigonometry

Degrees by default — pass mode="rad" for radians.

sin(angle, mode="deg") · cos(...) · tan(...)

The three primary trigonometric ratios.

sin(30)
# 0.5

cos(60)
# 0.5

tan(45)
# 1.0

cosec(angle, mode="deg") · sec(...) · cot(...)

The three reciprocal trigonometric ratios. Raise a clear error at an undefined point.

cosec(30)
# 2.0

sec(60)
# 2.0

cot(45)
# 1.0

Units

One function covers ten categories — length, mass, volume, area, time, speed, data, energy, pressure, and temperature. convert() figures out the category from the units themselves, so you never specify it; it just refuses to convert across categories (length to mass, say) with a clear error.

convert(value, from_unit, to_unit)

Converts between units within the same category. Case-insensitive, with generous aliases (km / kilometer / kilometers). Raises a clear error for incompatible categories.

convert(10, "km", "miles")
# 6.213711922373339

convert(98.6, "F", "C")
# 37.0

convert(1, "gb", "mb")
# 1024.0

Casting & Rounding

Type conversion and rounding, fixing two classic Python gotchas.

cast(value, target)

Casts a value — or an entire DataFrame column — to another type. Fixes Python's bool("False") gotcha (which is actually True!) and parses string literals like "[1, 2, 3]" into real lists instead of character-splitting.

cast("123", "int")
# 123

cast("False", "bool")
# False   (Python's own bool("False") is True!)

cast("[1, 2, 3]", "list")
# [1, 2, 3]

c["age"] = cast(c["age"], "int")   # casts a whole column

round(value, places=0)

Rounds using standard round-half-up — the way Excel works — instead of Python's "banker's rounding," where round(2.5) is actually 2.

round(2.5)
# 3   (Python's own round(2.5) gives 2)

round(1250, -2)
# 1300   (Python's own round(1250, -2) gives 1200)

Data Operations

The typical flow: load() a file into a variable, inspect it with loadh()/loadc()/loads(), optionally clean() it, then query it with sql() — real SQL, powered by DuckDB, run directly against your DataFrame by its variable name.

load(path)

Loads a CSV, Excel file, or Google Sheet into a DataFrame and previews it.

c = load("data.csv")
c = load("data.xlsx")
c = load("https://docs.google.com/spreadsheets/d/.../edit")

loadh(df) · loadc(df) · loads(df)

Preview just the column headers, the record count, or the schema (columns + datatypes).

loadh(c)
loadc(c)
loads(c)

sql(query)

Runs any SQL query against DataFrames already loaded in your script, by variable name. Powered by DuckDB — full SQL, including JOINs and window functions.

d = sql("select * from c where age > 30 order by age desc")

xtract(df, filename=None)

Exports a DataFrame to a CSV file in the current directory.

xtract(d)
xtract(sql("select * from c"), filename="filtered.csv")

clean(df)

Standardizes column names, strips whitespace from text columns, drops exact duplicates, and drops fully-empty rows. Returns a new DataFrame.

c = clean(c)

Lookups & Excel Functions

If you already know these from spreadsheets, they behave the same way here. Every function auto-detects your loaded dataset, so once you've called load(), you don't need to pass it in again — unless you've loaded more than one dataset, in which case pass it explicitly as df=.

vlook(value, return_col, df=None, default=...)

Searches the FIRST column (like Excel's VLOOKUP) and returns the matching row's value. Auto-detects your loaded dataset.

c = load("employees.csv")
vlook("Alice", "salary")
# 85000

vlook("Zoe", "salary", default=0)
# 0  (no error, since a fallback was given)

xlook(value, lookup_col, return_col, df=None, default=...)

Searches ANY column you specify (like Excel's XLOOKUP) — not limited to the first.

xlook("Engineering", "department", "name")
# "Alice"

countif(col, value, df=None)

Counts rows where col equals value.

countif("department", "Engineering")
# 3

sumif(col, value, sum_col, df=None)

Sums sum_col for rows where col equals value.

sumif("department", "Engineering", "salary")
# 250000

avgif(col, value, avg_col, df=None)

Averages avg_col for rows where col equals value.

avgif("department", "Engineering", "salary")
# 83333.33

AI

Bring your own API key from Anthropic, OpenAI, Gemini, or any OpenAI-compatible provider (Groq, Mistral, etc.) — the provider is detected automatically from the key's format, so you rarely need to specify it. Nothing here works without a key; everything else in Fyrefly does.

ask(data, question, api_key=None, provider=None, model=None, base_url=None)

Converts a plain-English question into SQL and runs it. Provider auto-detected from your key's format (Anthropic, OpenAI, Gemini, or Groq/any OpenAI-compatible API). Pass a tuple of DataFrames for joins across any number of datasets. Transient provider errors retry automatically.

c = load("sales.csv")

ask(c, "total sales by region last quarter?", api_key="AIza...")

ask((c, d), "top customers by total sales?", api_key="...")

insights(df, api_key=None, provider=None, model=None, base_url=None)

Generates a short, plain-English summary of a dataset — trends, ranges, and data-quality issues.

insights(c, api_key="AIza...")

Visualization

No API key, no separate charting library to learn. Call viz() directly with a kind, or use the shorter viz.bar(), viz.scatter(), etc. once you've loaded a single dataset — both produce the same chart.

viz(df, x=None, y=None, kind="auto", hue=None, title=None, figsize=(8,5), save=None)

Visualizes a DataFrame. kind="auto" picks a sensible chart based on your columns' types — heatmap, bar, histogram, or scatter.

viz(c)                               # auto -> correlation heatmap
viz(c, "age", "salary")              # auto -> scatter
viz(c, "department", "salary", kind="box")

viz.bar / .scatter / .line / .box / .violin / .hist / .kde / .pie / .heatmap / .pairplot

Namespace shortcuts for each chart type. Auto-detects your loaded DataFrame when exactly one is in scope.

c = load("data.csv")

viz.bar("department", "salary")
viz.scatter("age", "salary")
viz.heatmap()
viz.pairplot(hue="department")