Purpose
Synthetic data is often treated as clean by default, even when it can memorize source data, launder licensing problems, or contaminate evals.
A lineage checker that documents generation prompts, source classes, privacy tests, train/eval separation, and release constraints.
What it does
Validates a domain-specific AI governance packet, scores readiness, and returns concrete findings that contributors can improve.
Why it matters
AI systems are moving from chat into action. This project makes one hard operational risk easier to inspect, test, and govern in public.
Who should use it
Trace synthetic data back to generation policy, source class, and contamination risk. Builders can start with the CLI, then add adapters, fixtures, schemas, and integrations.
Quick Start
PYTHONPATH=src python3 -m unittest discover -s tests
python3 -m synthetic_data_lineage.cli sample
Example Packet
{
"dataset": {
"name": "synthetic-tickets",
"records": 10000
},
"generation": {
"model": "local-llm",
"promptPolicy": "no_pii"
},
"privacy": {
"nearestNeighborChecked": true,
"evalOverlapChecked": false
}
}
Contribution Tracks
Good first issues
- membership inference plugins
- benchmark contamination checks
- DVC integration
- privacy report templates
Core improvements
- Add JSON Schema validation.
- Add more real-world, non-sensitive fixtures.
- Improve scoring transparency and edge-case tests.
Integration work
- Build adapters for common AI frameworks.
- Add CI checks and report exports.
- Connect the packet format to operational workflows.