Data tool · CLI
WRANG
WRANG loads, inspects, cleans, transforms and exports datasets from the terminal with no boilerplate, interactively, from a script or in CI.
- Type
- Python package and CLI
- Stack
- Python 3.10+ · Polars · Rich · DuckDB (optional)
- Install
- pip install wrang
- License
- MIT
Overview
WRANG is a data-wrangling toolkit with three ways in: an interactive, menu-driven terminal session; non-interactive flags for scripts and CI; and a Python library whose modules can be imported on their own. It reads and writes CSV, Excel, Parquet and JSON or JSON Lines.
Where it came from
WRANG succeeds RIDE-CLI, the command-line half of my master’s project. Installing it adds both a wrang and a ride command, identical aliases, so existing ride scripts keep working.
Architecture
- Core modules: loader, inspector, explorer, cleaner, transformer and validator, each built on Polars DataFrames, with lazy scanning for large files.
- CLI layer: an interface, menus and formatters that render tables and panels with Rich.
- Optional extras:
wrang[viz]adds matplotlib and seaborn, andwrang[advanced]adds DuckDB and ConnectorX. - Configuration: user preferences such as the outlier factor and chunk size are kept in
~/.wrang/config.json.
Implementation
- Chainable cleaning and transforms.
DataCleanerandDataTransformerreturn themselves, so a pipeline reads top to bottom. There are nine imputation strategies, including distribution sampling and KNN. - SQL on any file.
--sqlruns DuckDB against the loaded file as a table nameddata. - Built for automation.
--inspect --output-format jsonprints a machine-readable overview.--chunk-sizestreams large files, and--comparediffs two datasets. - Data contracts. A schema can be inferred from data, saved as JSON or YAML, and validated later, with violations reported by severity.
- Reports.
--profilewrites a self-contained HTML profile of a dataset.
In use
Real output from wrang 0.2.2, installed from PyPI and run on the Titanic dataset bundled with the repository:
$ wrang titanic.csv --sql "SELECT Pclass, COUNT(*) AS n,
ROUND(AVG(Survived),3) AS survival_rate
FROM data GROUP BY Pclass ORDER BY Pclass"
Pclass n survival_rate
──────────────────────────────
1 216 0.63
2 184 0.473
3 491 0.242
3 row(s)
The repository’s README reports a test suite of 217 passing tests (plus 2 expected failures), and releases are published to PyPI from GitHub Actions.