Pandas GroupBy with a Custom Aggregation Function
Problem Given a pandas DataFrame, group rows by a key column and apply a custom (non-built-in) aggregation function to each group.
Be ready to discuss
- Forming groups with DataFrame.groupby(key), and what the resulting GroupBy object actually is — a lazy mapping from group key to row positions; nothing is computed until you aggregate.
- Choosing .agg() vs .apply(): groupby('category')['value'].agg(custom_function) reduces each group's column to a scalar, while groupby('category').apply(lambda g: custom_function(g)) hands you the whole sub-frame and may return a scalar, Series, or DataFrame.
- The other two members of the family: .transform() broadcasts a per-group result back to the original row shape, and .filter() drops whole groups by predicate.
- Performance: .apply() invokes a Python callback once per group, so it is slow when there are many small groups; prefer built-in Cython aggregations or decompose the custom logic into vectorized primitives.
- Multiple aggregations at once: passing a list or dict to .agg(), named aggregation, and dealing with the resulting MultiIndex columns.
- Gotchas: as_index=False vs. reset_index(), NaN group keys silently dropped unless dropna=False, and .apply() evaluating the first group twice for dtype inference.
asked …