Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

API reference

microdf exposes two classes. MicroSeries is a pandas.Series carrying a weight vector; MicroDataFrame is a pandas.DataFrame carrying a weight column. Both behave like their pandas counterparts, and the methods below either add a weighted estimator or preserve weights through an operation that would otherwise drop them.

import microdf as mdf

df = mdf.MicroDataFrame({"income": [10_000, 30_000, 120_000]}, weights=[800, 1_200, 50])
df.income.gini()

MicroSeries

Weighted aggregation

These have the same names as their pandas equivalents and return weighted results.

MethodSignatureDescription
sum(axis: Union[int, str, NoneType] = 0, skipna: bool = True, numeric_only: bool = False, min_count: int = 0, **kwargs) -> floatCalculates the weighted sum of the MicroSeries.
count(skipna: bool = True) -> floatCalculates the weighted count of the MicroSeries.
mean(skipna: bool = True) -> floatCalculates the weighted mean of the MicroSeries.
median(skipna: bool = True) -> floatCalculates the weighted median of the MicroSeries.
quantile(q: ndarray, skipna: bool = True) -> SeriesCalculates weighted quantiles of the MicroSeries.
var(ddof: int = 1, skipna: bool = True) -> floatCalculates the weighted variance of the MicroSeries.
std(ddof: int = 1, skipna: bool = True) -> floatCalculates the weighted standard deviation of the MicroSeries.
cov(other: Series, min_periods: Optional[int] = None, ddof: int = 1, *, skipna: bool = True) -> floatCalculate frequency-weighted covariance with another Series.
corr(other: Series, method: str = 'pearson', min_periods: Optional[int] = None, *, ddof: int = 1, skipna: bool = True) -> floatCalculate frequency-weighted Pearson correlation.
rank(pct: Optional[bool] = False) -> SeriesWeighted rank of each element.

Weight-preserving operations

Operations that change the shape or type of the data, overridden so weights stay aligned with their rows.

MethodSignatureDescription
groupby(*args, **kwargs) -> MicroSeriesGroupByGroup into MicroSeriesGroupBy, carrying weights into each group.
cumsum() -> SeriesCumulative sum of value times weight. Returns a plain pandas.Series: the weights have been applied and are not carried forward, so this is the one method here that does not preserve them.
astype(dtype, copy: Optional[bool] = True, errors: Optional[str] = 'raise') -> MicroSeriesConvert MicroSeries to specified data type while preserving weights.
clip(lower: Optional[float] = None, upper: Optional[float] = None, axis: Optional[int] = None, inplace: Optional[bool] = False, *args, **kwargs) -> MicroSeriesTrim values at the given thresholds, preserving weights.
round(decimals: Optional[int] = 0, *args, **kwargs) -> MicroSeriesRound each value, preserving weights.
repeat(repeats, axis=None)Repeat elements, repeating their weights alongside.
sqrt() -> MicroSeriesElement-wise square root, preserving weights.
copy(deep: Optional[bool] = True)Copy the series and its weights.
equals(other: MicroSeries) -> boolTrue when both the values and the weights are equal.
valuesattributeAccess underlying numpy array.
to_numpy(*args, **kwargs)Convert to numpy array.

Inequality and distribution

MethodSignatureDescription
gini(negatives: Optional[str] = None) -> floatCalculates Gini index.
top_1_pct_share() -> floatCalculates top 1% share.
top_10_pct_share() -> floatCalculates top 10% share.
top_50_pct_share() -> floatCalculates top 50% share.
bottom_50_pct_share() -> floatCalculates bottom 50% share.
top_0_1_pct_share() -> floatCalculates top 0.1% share.
top_x_pct_share(top_x_pct: float) -> floatCalculates top x% share.
bottom_x_pct_share(bottom_x_pct: float) -> floatCalculates bottom x% share.
t10_b50() -> floatCalculates ratio between the top 10% and bottom 50% shares.

Ranking

MethodSignatureDescription
decile_rank(negatives_in_zero: Optional[bool] = False)Calculate decile ranks (1-10) with optional zero decile for negatives.
quintile_rank() -> MicroSeriesCalculate weighted quintile ranks (1-5).
quartile_rank() -> MicroSeriesCalculate weighted quartile ranks (1-4).
percentile_rank() -> MicroSeriesCalculate weighted percentile ranks (1-100).

Variance from replicate weights

MethodSignatureDescription
replicate_standard_error(statistic: Callable, replicate_weights, method: str = 'jackknife', fay_k: Optional[float] = None, *, centering: str = 'full-sample') -> floatStandard error of statistic from a set of replicate weights.

replicate_standard_error accepts method of jackknife, brr, bootstrap, successive-difference, or fay (which also requires fay_k). Because it resamples rather than applying an analytic formula, it works for any statistic the series can compute, including the Gini coefficient and quantiles.

Weights

MethodSignatureDescription
set_weights(weights: ndarray, preserve_old: Optional[bool] = False) -> NoneSets the weight values.
nullify_weights() -> NoneSet all weights to 1, effectively making the Series unweighted.
weight() -> SeriesCalculates the weighted value of the MicroSeries.

MicroDataFrame

Weighted aggregation

MethodSignatureDescription
sum(axis: Union[int, str, NoneType] = 0, skipna: bool = True, numeric_only: bool = False, min_count: int = 0, **kwargs) -> Union[Series, MicroSeries, float]Sum numeric columns, weighting reductions across observations.
cov(min_periods: Optional[int] = None, ddof: int = 1, numeric_only: bool = False) -> DataFramePairwise frequency-weighted covariance of the columns.
corr(method: str = 'pearson', min_periods: int = 1, numeric_only: bool = False) -> DataFramePairwise frequency-weighted Pearson correlation of the columns.

Weight-preserving operations

MethodSignatureDescription
groupby(by: Union[str, list], *args, **kwargs) -> MicroDataFrameGroupByReturns a GroupBy object with MicroSeriesGroupBy objects for each column.
merge(right, how='inner', on=None, left_on=None, right_on=None, left_index=False, right_index=False, sort=False, suffixes=('_x', '_y'), copy=True, indicator=False, validate=None)Database-style join that carries the weight column through.
reset_index(level: Optional[int] = None, drop: Optional[bool] = False, inplace: Optional[bool] = False, col_level: Optional[int] = 0, col_fill: Optional[str] = '', allow_duplicates: Optional[bool] = None, names: Optional[list[str]] = None) -> Optional[MicroDataFrame]Reset the index, keeping weights aligned to their rows.
drop(labels=None, axis=0, index=None, columns=None, level=None, inplace=False, errors='raise')Drop rows or columns, keeping weights aligned to the remaining rows.
astype(dtype, copy: Optional[bool] = True, errors: Optional[str] = 'raise') -> MicroDataFrameConvert MicroDataFrame to specified data type while preserving weights.
copy(deep: Optional[bool] = True) -> MicroDataFrameCopy the frame and its weights.
equals(other: MicroDataFrame) -> boolTrue when both the values and the weights are equal.

Poverty

MethodSignatureDescription
poverty_rate(income: str, threshold: str) -> floatCalculate poverty rate, i.e., the population share with income below their poverty threshold.
poverty_gap(income: str, threshold: str) -> floatCalculate poverty gap, i.e., the total gap between income and poverty thresholds for all people in poverty.
poverty_count(income: Union[MicroSeries, str], threshold: Union[MicroSeries, str]) -> intCalculates the number of entities with income below a poverty threshold.
deep_poverty_rate(income: str, threshold: str) -> floatCalculate deep poverty rate, i.e., the population share with income below half their poverty threshold.
deep_poverty_gap(income: str, threshold: str) -> floatCalculate deep poverty gap, i.e., the total gap between income and half of poverty thresholds for all people in deep poverty.
squared_poverty_gap(income: str, threshold: str) -> floatCalculate squared poverty gap, i.e., the total squared gap between income and poverty thresholds for all people in poverty. Also known as the poverty severity index.

Weights

MethodSignatureDescription
set_weights(weights: Union[ndarray, str], preserve_old: Optional[bool] = False) -> NoneSets the weights for the MicroDataFrame.
set_weight_col(column: str, preserve_old: Optional[bool] = False) -> NoneSets the weights for the MicroDataFrame by specifying the name of the weight column.
nullify_weights() -> NoneSet all weights to 1, effectively making the DataFrame unweighted.

Module-level functions

FunctionDescription
microdf.replicate_varianceVariance of a statistic from replicate weights.
microdf.replicate_standard_errorSquare root of the above.

A note on estimator conventions

Quantiles follow the inverse cumulative distribution function, so results can be checked against survey::svyquantile in R. Weighted variance treats weights as frequency weights, so integer weights agree with numpy computed on the replicated sample. Top-share cutoffs split a record that straddles the boundary in proportion, rather than assigning it wholly to one side.