microdf exposes two classes. MicroSeries is a pandas.Series carrying a weight
vector; MicroDataFrame is a pandas.DataFrame carrying a weight column. Both
behave like their pandas counterparts, and the methods below either add a
weighted estimator or preserve weights through an operation that would otherwise
drop them.
Operations that change the shape or type of the data, overridden so weights stay aligned with their rows.
Method
Signature
Description
groupby
(*args, **kwargs) -> MicroSeriesGroupBy
Group into MicroSeriesGroupBy, carrying weights into each group.
cumsum
() -> Series
Cumulative sum of value times weight. Returns a plain pandas.Series: the weights have been applied and are not carried forward, so this is the one method here that does not preserve them.
Standard error of statistic from a set of replicate weights.
replicate_standard_error accepts method of jackknife, brr, bootstrap,
successive-difference, or fay (which also requires fay_k). Because it
resamples rather than applying an analytic formula, it works for any statistic
the series can compute, including the Gini coefficient and quantiles.
Calculate poverty rate, i.e., the population share with income below their poverty threshold.
poverty_gap
(income: str, threshold: str) -> float
Calculate poverty gap, i.e., the total gap between income and poverty thresholds for all people in poverty.
poverty_count
(income: Union[MicroSeries, str], threshold: Union[MicroSeries, str]) -> int
Calculates the number of entities with income below a poverty threshold.
deep_poverty_rate
(income: str, threshold: str) -> float
Calculate deep poverty rate, i.e., the population share with income below half their poverty threshold.
deep_poverty_gap
(income: str, threshold: str) -> float
Calculate deep poverty gap, i.e., the total gap between income and half of poverty thresholds for all people in deep poverty.
squared_poverty_gap
(income: str, threshold: str) -> float
Calculate squared poverty gap, i.e., the total squared gap between income and poverty thresholds for all people in poverty. Also known as the poverty severity index.
Quantiles follow the inverse cumulative distribution function, so results can be
checked against survey::svyquantile in R. Weighted variance treats weights as
frequency weights, so integer weights agree with numpy computed on the
replicated sample. Top-share cutoffs split a record that straddles the boundary
in proportion, rather than assigning it wholly to one side.