Dataset
A dataset is a collection of examples organized for a task, such as text, images, labels, or interactions. Training, validation, and test data serve different purposes. A retrieval corpus supplies evidence at answer time in RAG.
Separate the purposes
| Split | Purpose | Avoid |
|---|---|---|
| Training | Fit parameters and preprocessing | Including test questions and answers |
| Validation | Choose configurations | Treating repeatedly tuned scores as final evidence |
| Test | Assess a frozen system | Using the result to guide further tuning |
There is no universal split ratio. Our suggested starting point is to identify the deployment scenario and choose an entity, document-source, or time-based split. See data leakage.
Describe the data
A dataset card can document provenance, collection, language, size, label definitions, permissions, intended uses, and known biases. Report examples and tokens separately, with the version and the stage at which counts were measured.
Example: support questions
This invented record illustrates a possible schema:
{"id":"example-001","group_id":"ticket-demo","question":"How do I reset my password?","answer":"Open account settings and verify your identity.","source":"help/password","version":"v1","split":"train"}
We would keep rewrites from one ticket in the same split. A retrieval corpus could additionally retain source locations and access metadata; a training dataset needs a separate assessment of label quality and training suitability.
Sources
- Hugging Face: Dataset Cards — dataset documentation and metadata.
- Datasheets for Datasets — documenting provenance, composition, and uses.
- scikit-learn: Cross-validation — grouped splits.