Data tool · CLI

WRANG

WRANG loads, inspects, cleans, transforms and exports datasets from the terminal with no boilerplate, interactively, from a script or in CI.

Type
Python package and CLI
Stack
Python 3.10+ · Polars · Rich · DuckDB (optional)
Install
pip install wrang
License
MIT

Overview

WRANG is a data-wrangling toolkit with three ways in: an interactive, menu-driven terminal session; non-interactive flags for scripts and CI; and a Python library whose modules can be imported on their own. It reads and writes CSV, Excel, Parquet and JSON or JSON Lines.

Where it came from

WRANG succeeds RIDE-CLI, the command-line half of my master’s project. Installing it adds both a wrang and a ride command, identical aliases, so existing ride scripts keep working.

Architecture

  • Core modules: loader, inspector, explorer, cleaner, transformer and validator, each built on Polars DataFrames, with lazy scanning for large files.
  • CLI layer: an interface, menus and formatters that render tables and panels with Rich.
  • Optional extras: wrang[viz] adds matplotlib and seaborn, and wrang[advanced] adds DuckDB and ConnectorX.
  • Configuration: user preferences such as the outlier factor and chunk size are kept in ~/.wrang/config.json.

Implementation

  • Chainable cleaning and transforms. DataCleaner and DataTransformer return themselves, so a pipeline reads top to bottom. There are nine imputation strategies, including distribution sampling and KNN.
  • SQL on any file. --sql runs DuckDB against the loaded file as a table named data.
  • Built for automation. --inspect --output-format json prints a machine-readable overview. --chunk-size streams large files, and --compare diffs two datasets.
  • Data contracts. A schema can be inferred from data, saved as JSON or YAML, and validated later, with violations reported by severity.
  • Reports. --profile writes a self-contained HTML profile of a dataset.

In use

Real output from wrang 0.2.2, installed from PyPI and run on the Titanic dataset bundled with the repository:

$ wrang titanic.csv --sql "SELECT Pclass, COUNT(*) AS n,
    ROUND(AVG(Survived),3) AS survival_rate
    FROM data GROUP BY Pclass ORDER BY Pclass"

  Pclass   n     survival_rate
 ──────────────────────────────
  1        216   0.63
  2        184   0.473
  3        491   0.242

3 row(s)

The repository’s README reports a test suite of 217 passing tests (plus 2 expected failures), and releases are published to PyPI from GitHub Actions.

Contact

Have something interesting in mind?

I’m always interested in thoughtful collaborations, interesting research problems, and opportunities to build useful things.