> For the complete documentation index, see [llms.txt](https://docs.sdv.dev/rdt/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.sdv.dev/rdt/rdt-reversible-data-transforms.md).

# RDT: Reversible Data Transforms

How much effort are you spending in cleaning and processing your data?

RDT (Reversible Data Transforms) is a [Python library](https://github.com/sdv-dev/RDT) that translates between real world data and cleaned, numerical data that's ready for data science.

![](https://2225246359-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FVGX92M819eIp0rMg5elc%2Fuploads%2FUTGsaoq8ummfcIFB72i9%2Frdt_hypertransformer-transforming-into-numerical-data-and-back_June%2003%202025.png?alt=media\&token=4ed03e04-48d8-45f3-84f1-0a558e38948b)

## More than a formatting library

Cleaning and formatting raw data is a foundational element of RDT. But you can use the library to do much more.

### :1234: Statistical Processing

Normalize your data using statistical processes. This is especially useful for data science and machine learning projects.

![](https://2225246359-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FVGX92M819eIp0rMg5elc%2Fuploads%2FkFEz1qGsTm53Y2ZoiNCQ%2Frdt_statistical_processing.png?alt=media\&token=95b5ee01-41c6-4d42-be9e-f66266b3fba6)

### :unlock: Anonymizing Sensitive Data

Protect sensitive data while preserving the overall data format. Using RDTs, you can remove and anonymize Personal Identifiable Information. Use it to generate random, fake values that look like the original ones.

![](https://2225246359-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FVGX92M819eIp0rMg5elc%2Fuploads%2FUADe4W76xAf5iO2im0NE%2Frdt_reversible-data-transforms-anonymizing-sensitive-data_June%2002%202025.png?alt=media\&token=2e604799-2797-4a35-b147-eb4b08a9dd33)

### :gem: Extracting Deeper Meaning

Licensed users can extract deeper concepts that are embedded inside the data. This is particularly useful for complex data types that have a rich, real-world meaning.

![](https://2225246359-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FVGX92M819eIp0rMg5elc%2Fuploads%2FuO3UPPXnoOEJyXijFtuC%2Frdt_reversible-data-transforms-extracting-deeper-meaning_June%2002%202025.png?alt=media\&token=8a7cfb88-6559-4f60-8f0a-a1319a93604d)

## Use Cases

We created RDTs with the goal of **generating synthetic data**. The RDT library transforms the raw data for machine learning, and then reverse transforms machine-generated data to match the original. Synthetic data remains a top use case for RDT today.

{% hint style="info" %}
If you'd like to use RDT for synthetic data, we recommend installing the [sdv library](https://sdv.dev/SDV/getting_started/install.html). It will automatically download RDT, along with other libraries to support synthetic data generation & evaluation.&#x20;
{% endhint %}

RDT can be useful beyond the synthetic data space. You can use RDT for statistical preprocessing, contextual anonymization, or adding differential privacy as standalone projects. For more information, see [Use Cases](/rdt/resources/use-cases.md).&#x20;

## Owned & Maintained by DataCebo

The RDT library is a part of the [Synthetic Data Vault Project](https://sdv.dev/), first created at MIT's [Data to AI Lab](http://dai.lids.mit.edu/) in 2016. After 4 years of research and traction with enterprise, we created DataCebo in 2020 with the goal of growing the project.

Today, [DataCebo](https://datacebo.com/) is the proud developer of the SDV, the largest ecosystem for synthetic data generation & evaluation.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.sdv.dev/rdt/rdt-reversible-data-transforms.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
