Skip to content
Open
Show file tree
Hide file tree
Changes from 6 commits
Commits
Show all changes
100 commits
Select commit Hold shift + click to select a range
3cedc86
Merge pull request #2 from bids-standard/master
ericearl May 20, 2025
11fbb47
Merge pull request #3 from bids-standard/master
ericearl May 30, 2025
0ef9fdf
[ENH] Integrate BEP036 - Phenotypic Data Guidelines
surchs May 30, 2025
0a640e6
Update phenotype.md and data-summary-files.md
ericearl May 30, 2025
a19512b
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] May 30, 2025
5718888
Update data-summary-files.md and phenotypic-and-assessment-data.md
ericearl May 30, 2025
8f54e94
Apply suggestions from code review
ericearl May 30, 2025
94cb476
Apply suggestions from code review
ericearl May 30, 2025
142c460
Apply suggestions from code review
ericearl May 30, 2025
8b78359
Update src/modality-agnostic-files/phenotypic-and-assessment-data.md
ericearl May 30, 2025
60f712a
Update mkdocs.yml
ericearl May 30, 2025
e62b5cc
Update src/modality-agnostic-files/phenotypic-and-assessment-data.md
ericearl May 30, 2025
ac097aa
Update phenotype.md to have a macro table from schema
ericearl Jun 24, 2025
32fedd0
Update src/schema/rules/tabular_data/modality_agnostic.yaml
ericearl Jun 24, 2025
aacda9b
Update modality_agnostic.yaml
ericearl Jun 24, 2025
fd5ff2d
Update phenotype.md appendix and modality_agnostic.yaml schema
ericearl Jun 24, 2025
dd65b5e
Update modality_agnsotic.yaml
ericearl Jun 24, 2025
f4205e8
add missing column objects, use existing acq column definition (#4)
rwblair Jul 17, 2025
0eba71d
Updates for BEP036
ericearl Jul 17, 2025
d3631a8
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 17, 2025
f4939ad
Updates phentoype.md and Guideline 3 in the modality agmpstic section
ericearl Jul 17, 2025
abd5c2b
Update modality_agnostic.yaml
ericearl Jul 17, 2025
ec2c53d
Update phenotypic-and-assessment_data.md
ericearl Jul 17, 2025
7639001
Merge branch 'master' into master
ericearl Jul 17, 2025
8b38859
Merge branch 'bids-standard:master' into master
ericearl Sep 15, 2025
d1141a0
Updates to remove demographics file and add AdditionalValidation field
ericearl Sep 17, 2025
6c6ee8b
Attempting to satisfy the CI and remark
ericearl Sep 17, 2025
9f8afec
Merge branch 'bids-standard:master' into master
ericearl Sep 17, 2025
8fa89bc
Update src/schema/rules/tabular_data/modality_agnostic.yaml
effigies Sep 18, 2025
ff86669
Update columns.yaml
ericearl Sep 18, 2025
ede68ef
Update modality_agnostic.yaml
ericearl Sep 18, 2025
e8ab5dd
Update phenotype.md appendix examples, participants schema, and other…
ericearl Sep 22, 2025
3490e9d
Update modality_agnostic.yaml
ericearl Sep 23, 2025
d02e0bf
Update modality_agnostic.yaml
ericearl Sep 23, 2025
f8d6333
Update modality_agnostic.yaml
ericearl Sep 23, 2025
bd083c0
Update phenotype appendix
ericearl Sep 24, 2025
ec2703b
Update phenotype appendix
ericearl Sep 24, 2025
00d8f25
Delete .vscode/settings.json
ericearl Sep 25, 2025
41f0f70
Apply suggestions from code review
ericearl Sep 25, 2025
cdfc0d2
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Sep 25, 2025
6cbb4ee
Update BEP036 files more
ericearl Sep 29, 2025
fe3ddab
Update appendices/phenotype.md
ericearl Sep 29, 2025
40f6751
Update src/schema/rules/tabular_data/modality_agnostic.yaml
ericearl Sep 30, 2025
2fd12d7
Update src/schema/objects/columns.yaml
ericearl Sep 30, 2025
a0cab8b
Update src/appendices/phenotype.md
ericearl Sep 30, 2025
c0bd78a
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Sep 30, 2025
97917f0
Update src/modality-agnostic-files/data-summary-files.md
ericearl Sep 30, 2025
80683f6
Merge branch 'master' into master
surchs Sep 30, 2025
76932fe
Move longitudinal age section to point 1
surchs Sep 30, 2025
32f994e
Revise section 5
surchs Sep 30, 2025
3a602ca
fix: Participant ID mismatch check
effigies Oct 1, 2025
f8d492e
fix(schema): Resolve a couple issues
effigies Oct 1, 2025
7f1eb09
feat: Expanded validation checks
effigies Oct 1, 2025
5c55eb9
deduplicate
effigies Oct 1, 2025
c0951f3
Promote Participants
effigies Oct 1, 2025
69183c7
feat: Make MeasurementToolMetadata recommended for phenotype files
effigies Oct 2, 2025
d3f1d0d
Updates for community review of BEP036
ericearl Oct 11, 2025
a9e2d4b
Merge branch 'master' into master
ericearl Oct 11, 2025
b60eac1
Update phenotype.md and data-summary-files.md
ericearl Nov 4, 2025
b95feec
Merge branch 'bids-standard:master' into master
ericearl Nov 4, 2025
d8b34f3
Update phenotype.md and phenotypic-and-assessment-data.md
ericearl Nov 4, 2025
b908a19
Update src/appendices/phenotype.md
ericearl Jan 21, 2026
c156df4
Catch up to the latest spec (1/21/26) (#5)
ericearl Jan 21, 2026
d2a67f6
Update modality_agnostic.yaml schema file
ericearl Jan 21, 2026
b19356f
Merge branch 'master' into master
effigies Jan 22, 2026
4be4556
Update phenotype.md and modality_agnostic.yaml
ericearl Jan 23, 2026
db34902
Merge branch 'master' into bep036-patch-mkdocs-macro
ericearl Jan 28, 2026
a7abb76
Merge remote-tracking branch 'upstream/master' into bep036-patch-mkdo…
effigies Mar 23, 2026
9dc6637
feat(schema): Add glob and zip functions to expression language
effigies Mar 6, 2026
37df2e4
feat(schema): Add dataset.{participants,sessions}_tsv to context
effigies Mar 6, 2026
878bc4a
rf(schema): Rewrite rules to use new context, glob()
effigies Mar 6, 2026
0708e30
rf(schema): Remove unused dataset.subjects from context
effigies Mar 6, 2026
4b3fdef
test(schema): Add expression tests
effigies Mar 6, 2026
46f2777
test(bst): Drop Subjects class from context types
effigies Mar 6, 2026
69c8cde
Merge remote-tracking branch 'origin/rf/tables' into bep036-patch-mkd…
effigies Mar 23, 2026
882734a
feat(schema): Add check for session_id columns
effigies Mar 17, 2026
abf16ca
feat(schema): Require aggregated session files
effigies Mar 17, 2026
2ec8859
feat(schema): Require session directories if sessions.tsv is defined
effigies Mar 17, 2026
b050199
fix: Require subject entity to trigger session entity check
effigies Mar 18, 2026
512c8d4
feat: Check for sessions.tsv file on any file with session entity
effigies Mar 18, 2026
1c9fef6
feat(metaschema): Allow for more complete issues
effigies Mar 18, 2026
ecda3ad
feat(schema): add BEP036 phenotype check rules
effigies Mar 20, 2026
4669da3
fix: Limit PhenotypeSessionAggregation to Phenotype validation
effigies Mar 23, 2026
aefcb5d
Update columns.yaml
ericearl Apr 23, 2026
d45ce4c
Update to latest BIDS spec (#7)
ericearl Jun 11, 2026
273b88e
BEP036 IndexColumns rewrite draft (#8)
ericearl Jun 18, 2026
5dce352
Update phenotype.md and longitudinal-and-multi-site-studies.md
ericearl Jun 18, 2026
12e7db1
Merge branch 'master' into bep036-patch-mkdocs-macro
ericearl Jun 18, 2026
86f8ecb
Update phenotype.md
ericearl Jun 18, 2026
b7cc269
yml lint fixes
rwblair Jul 2, 2026
ebab3da
Merge branch 'master' of github.com:bids-standard/bids-specification …
rwblair Jul 2, 2026
4d581e2
update extended phenotype checks to test for tool-demographics file i…
rwblair Jul 7, 2026
b20711c
folder -> directory
rwblair Jul 8, 2026
5a14754
Merge branch 'master' into bep036-patch-mkdocs-macro
effigies Jul 17, 2026
e84dac1
Consolidate default sidecar table
effigies Jul 17, 2026
bcade15
Updates to address Sam's and Ross' comments
ericearl Jul 17, 2026
bafec92
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 17, 2026
2194f61
Merge branch 'master' into bep036-patch-mkdocs-macro
effigies Jul 19, 2026
9bc39d3
Merge branch 'master' into bep036-patch-mkdocs-macro
effigies Jul 21, 2026
df8b4a7
Update phenotype.md, data-summary-files.md, and modality_agnostic.yaml
ericearl Jul 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
331 changes: 331 additions & 0 deletions src/appendices/phenotype.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,331 @@
# Tabular phenotypic data guidelines

This appendix is a collection of guidelines and examples for creating well-organized aggregated tabular phenotypic data.

## Guidelines

These guidelines are all **RECOMMENDED** when preparing
tabular phenotypic data like the
participants file, sessions file, demographics file,
or phenotypic and assessment data.
The language below uses REQUIRED, MUST, and others to imply
these are the requirements for these **RECOMMENDED** guidelines.

### 1. Always pair tabular data with data dictionaries

Tabular phenotypic data MUST be prepared as one pair of a tabular file
in tab-separated value (TSV) format and a corresponding data dictionary
in JavaScript Object Notation (JSON) format.
Comment thread
ericearl marked this conversation as resolved.
Outdated

### 2. Aggregate data across sessions

Aggregation refers to the contents of the TSV file. It is REQUIRED
to collect all participant data into one TSV per tabular phenotypic file.

### 3. Ensure minimal annotation for phenotypic and assessment data

In phenotypic and assessment data each measurement tool has an independent
aggregated data TSV file in which the user collects all subjects, sessions,
and/or runs of data as one entry per row (with a row defined by
the smallest unit of acquisition). In other words:

1. Each row MUST start with `participant_id`.
2. Each TSV file MUST contain a `session_id` column when
multiple [sessions](../glossary.md#session-entities)[^1] are present
in the data set regardless of whether those sessions are in
the `phenotype/` data, `sub-<label>/` data, or a combination of the two.
3. If more than one of the same measurement tool is acquired within
the same `session_id`, a `run` column MUST be added.
4. To encode the acquisition time for a measurement tool’s `session_id`,
add the `session_id` to the sessions file and
include the OPTIONAL `acq_time` column.
Comment thread
ericearl marked this conversation as resolved.
Outdated

To summarize this guideline as a table:

| **Column name** | **Requirement** | **Description** |
| :--------------- | :-------------- | :-------------- |
| `participant_id` | REQUIRED | MUST be the first column in the file. Note that data for one participant MAY be represented across multiple rows in case of multiple sessions or runs, and therefore the entry in the `participant_id` column will be repeated. |
| `session _id` | CONDITIONAL ; If sessions are defined in the dataset | A `session_id` column MUST be added to all tabular files in the phenotype directory as soon as multiple sessions are present in the data set regardless of whether those sessions are in the `phenotype/` data, `sub-<label>/` data, or a combination of the two. |
| `run` | CONDITIONAL ; If there are multiple runs within any session | A chronological `run` number is used when a measurement tool or assessment described by a tabular file was repeated within a session. |
| `acq_time` | OPTIONAL | If acquisition time is available, the `acq_time` column CAN be used to record the time of acquisition of each row in the tabular file. |

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@effigies Is there a macro we can use maybe to clean this table up? How would you recommend making this table more manageable?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Participants:
selectors:
- path == "/participants.tsv"
initial_columns:
- participant_id
columns:
participant_id:
level: required
description_addendum: |
There MUST be exactly one row for each participant.
species: recommended
age: recommended
sex: recommended
handedness: recommended
strain: recommended
strain_rrid: recommended
index_columns: [participant_id]
additional_columns: allowed

<!-- This block generates a columns table.
The definitions of these fields can be found in
src/schema/rules/tabular_data/*.yaml
and a guide for using macros can be found at
https://github.com/bids-standard/bids-specification/blob/master/macros_doc.md
-->
{{ MACROS___make_columns_table("modality_agnostic.Participants") }}


Furthermore, if you have to add a `session_id` column to the
tabular phenotypic data, you then MUST also introduce a session directory to the
imaging data, even if only one imaging session has been created.
This rule can be considered as "**if anyone uses sessions, everyone uses sessions**."
And vice versa, if imaging data has session directories,
all imaging data and tabular phenotypic data MUST have sessions.

This produces a file in which same-participant entries can take up as many rows
as needed according to the smallest unit of acquisition.
The combination of values in the `participant_id`, `session_id`, and `run` (if present)
columns MUST be unique for the entire tabular file.

### 4. Add `MeasurementToolMetadata` to each tabular phenotypic measurment tool
Comment thread
ericearl marked this conversation as resolved.
Outdated

Whenever possible, it is RECOMMENDED to add `MeasurementToolMetadata` to
each `phenotype/<measurement_tool_name>.json` data dictionary.
Comment thread
ericearl marked this conversation as resolved.
Outdated
This improves reusability and provides clarity about the measurement tool.

### 5. Use the demographics file for common variables about participants

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copying from https://github.com/surchs/bids-specification/pull/1/files#r2103117486

For this section, would it make sense to suggest that demo-like information be prioritized in this file rather than participants.tsv, making the latter primarily a list of subject IDs? I haven't seen this explicitly addressed anywhere, though I'm unsure if it's something we want to formalize 😬
Something like this could follow the paragraph?:

When all demographic data is stored in phenotype/demographics.tsv, participants.tsv may serve primarily as a minimal listing of subject identifiers with only the participant_id column.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree. It'd be good to mention this.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, I am new to commenting a PR and also relatively new to BIDS - so my comment might be stupid.

However, I want to draw attention to a potential issue with the handling repeated usage of a phenotype/ measurement. Currently, if a measurement is conducted more than once, a run_id column MUST be added that "corresponds to an existing run- entity used in a filename(s)" (https://bids-specification--2123.org.readthedocs.build/en/2123/modality-agnostic-files/phenotypic-and-assessment-data.html).

However, in my case (and I guess many others), a phenotype measurement tool might be conducted independently of a run in an experimental task. For example, I conduct a mood questionnaire 7 times over the course of one session in which multiple tasks (with multiple runs) are conducted. Thus, I think it is not reasonable that the run_id column in the phenotype file corresponds to an existing run_ in a (task) filename. Instead, something like a measurement_timepoint_id (or similar) that is not related to the run_ in other files might be a good idea. Again, I am sorry if my comment is stupid or I misunderstood something.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Along this line: The current proposal states:

A participant identifier of the form sub-, matching a participant entity found in the dataset. Note that data for one participant MAY be represented across multiple rows in case of multiple sessions or runs, and therefore the entry in the participant_id column will be repeated. The combination of participant_id, session_id and run_id MUST be unique.

The problem with this is that it doesn't cover the following cases which are in many datasets that I am familiar with:

Example 1: A dataset in which there are multiple training sessions on some equipment or on some test procedure on multiple different days before any imaging or any data that would have a run occurs.

Example 2: Sleep questionnaires and other questionnaires that administered multiple times un-associated with any imaging sessions.

Both of these require another column identifier (not associated with the run).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@VisLab Great to hear from you! Thanks for reading the BEP. Now to your comments...

Example 1: A dataset in which there are multiple training sessions on some equipment or on some test procedure on multiple different days before any imaging or any data that would have a run occurs.

We may not have made it explicit, but here is this example in the HTML preview of the appendix. It covers what to do when there's "1 participant with 2 sessions, where 1 session is only tabular phenotype and the other is only imaging". The same idea extends to more than 2 sessions of some combination of acquired data and tabular phenotypic data. For instance, in your Example 1, we could say:

Define many phenotypic session names

  • ses-day1equipmenttraining1
  • ses-day1equipmenttraining2
  • ses-day1equipmenttraining3
  • ses-day2equipmenttraining1
  • ses-day2equipmenttraining2
  • ses-preproceduretraining
  • ses-procedure
  • ses-postprocedure
  • ...

Store most, if not all, in one aggregated file

  • phenotype/equipment_training.tsv
    • Where participant_id, session_id, run_id, and day are joint indexes.
    • Yes, day is not a a validate-able column, but you could alternatively create a phenotype/day#_equipment_training.tsv if you want stricter validation support.
    • Remember that run_id is not about tying the data to a particular task's run, but instead to differentiate multiple runs of the same tabular phenotypic test, survey, assessment, questionnaire, or lab result within the same participant_id's session_id.

Example 2: Sleep questionnaires and other questionnaires that administered multiple times un-associated with any imaging sessions.

The answer to example 2 is the appendix example I linked above, and now here.

I hope this helps and let me know if you don't think this resolves your comment. Or if perhaps there's something else we should consider which you can propose?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@timdressler

I want to draw attention to a potential issue with the handling repeated usage of a phenotype/ measurement. Currently, if a measurement is conducted more than once, a run_id column MUST be added that "corresponds to an existing run- entity used in a filename(s)" (https://bids-specification--2123.org.readthedocs.build/en/2123/modality-agnostic-files/phenotypic-and-assessment-data.html).

Wowza, good catch! I didn't notice that incorrect set of words there. It currently says:

(run_id is) A run identifier that corresponds to an existing run- entity used in a filename(s). A chronological run number is used when a measurement tool or assessment described by a tabular file was repeated within a session.

What it should say:

(run_id is) A run identifier to differentiate multiple runs of the same measurement tool. A chronological run number is used when a measurement tool or assessment described by a tabular file is repeated within a session.

Thanks! I'll edit that.

For example, I conduct a mood questionnaire 7 times over the course of one session in which multiple tasks (with multiple runs) are conducted.

Does the text edit above of "what it should say" make better sense and resolve the run_id labeling issue?


Some studies collect demographics into their own tabular phenotypic data file already.
In these cases, it is RECOMMENDED to house this data in the `phenotype/` directory
as a TSV called `demographics.tsv` and its corresponding data dictionary JSON
called `demographics.json`.

### 6. Store longitudinal age in the demographics file

It is RECOMMENDED to use the `age` column to record participant age
at every session in longitudinal or multi-session data sets.
This reduces data duplication across tabular data files. The `Units` of `age`
do not have to be years so long as the units of the age
are written in `phenotype/demographics.json`.
Consider participant privacy or study objectives when selecting
the `Units` of `age` or the accuracy of `age` data.

### 7. Use the sessions file at the root level

If there is more than one session for any one participant, then
it is REQUIRED to provide a sessions file at the dataset root.
The sessions file MUST list all sessions for all subjects across
imaging and tabular phenotypic data.

When a sessions file is in use, you MUST NOT provide additional sessions files
at the participant-level which would otherwise use the inheritance principle.
Comment thread
ericearl marked this conversation as resolved.
Outdated
If a sessions file is provided, then it MUST begin with a `participant_id` column
followed immediately by a `session_id` column. The data dictionary JSON file’s
`session_id` field MUST include `Levels` with the description of each `session_id`.

### 8. Record acquisition time of sessions with `acq_time`

Whenever possible, it is RECOMMENDED to also collect acquisition time for
tabular phenotypic data and store the time of acquisition[^2] of each row
inside a column named `acq_time` in the sessions file.
This is consistent with how acquisition time is recorded for MRI data
and other time-sensitive measurements (for example systolic blood pressure).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with this, but I find it a bit confusingly worded.

store the time of acquisition[^2] of each row inside a column named acq_time in the sessions file.

Essentially what we're saying is: please record the acq_time for all sessions. And when you do, put that in the sessions.tsv


When needed to preserve participant privacy, you SHOULD record
relative acquisition times with respect to the earliest session.
Relative session acquisition times MAY be listed as durations from
the earliest session (baseline) in days, months, or years
using the `acq_time` column.

## Summary

This appendix described seven guidelines for best tabular phenotypic data.
A short summary table here describes when to use which files.

| File | Single session data | Multiple session data |
| :----------------------------- | :------------------ | :-------------------- |
| Participants | RECOMMENDED | RECOMMENDED |
| Phenotypic and assessment data | RECOMMENDED | RECOMMENDED |
| Sessions | OPTIONAL | REQUIRED |
| Demographics | OPTIONAL | RECOMMENDED |

## Examples

What follows are a few common use case examples for tabular phenotypic files.

### 1 participant session with both non-tabular and tabular phenotypic data

File tree

```Text
phenotype/
<measurement_tool_name>.json
<measurement_tool_name>.tsv
sub-01/anat/
sub-01_T1w.json
sub-01_T1w.nii.gz
```

Contents of `phenotype/<measurement_tool_name>.tsv`

```Text
participant_id measurement_1 measurement_2
sub-01 value1 value2
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

### 1 participant with 2 sessions, where 1 session is only tabular phenotype and the other is only imaging

With only one imaging and one phenotypic session each in this example you might want
to merge both imaging and phenotypic data under one session. But it is more correct to
have separate sessions for the imaging and phenotypic data, especially if
the sessions were collected days, weeks, or months apart. You can denote both sessions
and their acquisition time in the `sessions.tsv` file and have `session_id` `Levels` noted
in the `sessions.json` sidecar. Below are a CORRECT and an INCORRECT example
of prepared data following these guidelines.

#### CORRECT

File tree

```Text
phenotype/
<measurement_tool_name>.json
<measurement_tool_name>.tsv
sub-01/ses-MRI/anat/
sub-01_ses-MRI_T1w.json
sub-01_ses-MRI_T1w.nii.gz
```

Contents of `phenotype/<measurement_tool_name>.tsv`

```Text
participant_id session_id measurement_1 measurement_2
sub-01 ses-pheno value1 value2
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

#### INCORRECT

File tree

```Text
phenotype/
<measurement_tool_name>.json
<measurement_tool_name>.tsv
sub-01/anat/
sub-01_T1w.json
sub-01_T1w.nii.gz
```

Contents of `phenotype/<measurement_tool_name>.tsv`

```Text
participant_id measurement_1 measurement_2
sub-01 value1 value2
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

A session directory **MUST** be present in the participant directory and
the `session_id` column **MUST** be present in `<measurement_tool_name>.tsv` as well.
Sessions must be used consistently for the combination of tabular and
non-tabular phenotypic data.

### 2 participants with a mix of tabular phenotypic data and imaging sessions

File tree

```Text
phenotype/
<measurement_tool_name>.json
<measurement_tool_name>.tsv
sub-01/
ses-MRI1/
anat/
sub-01_ses-MRI1_T1w.json
sub-01_ses-MRI1_T1w.nii.gz
ses-MRI2/
anat/
sub-01_ses-MRI2_T1w.json
sub-01_ses-MRI2_T1w.nii.gz
sub-02/
ses-MRI1/
anat/
sub-02_ses-MRI1_T1w.json
sub-02_ses-MRI1_T1w.nii.gz
```

Contents of `phenotype/<measurement_tool_name>.tsv`

```Text
participant_id session_id measurement_1 measurement_2
sub-01 ses-pheno1 value1 value2
sub-02 ses-pheno1 value3 value4
sub-02 ses-pheno2 value5 value6
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

### 3 participants with 3 different kinds of sessions among them

The `ses-baseline` session collects an MRI and tabular phenotypic data.

File tree

```Text
participants.json
participants.tsv
sessions.json
sessions.tsv
phenotype/
demographics.json
demographics.tsv
...
sub-01/
ses-baseline/
ses-followupMRI/
sub-02/
ses-baseline/
sub-03/
ses-baseline/
ses-followupMRI/
```

Contents of `sessions.tsv`.

```Text
participant_id session_id acq_time
sub-01 ses-baseline 2001-01-01T12:05:00
sub-01 ses-followupMRI 2001-07-01T13:33:00
sub-01 ses-interview 2002-01-01T11:21:00
sub-02 ses-baseline 2001-04-01T11:01:00
sub-02 ses-interview 2002-04-01T14:08:00
sub-03 ses-baseline 2001-09-01T11:45:00
sub-03 ses-followupMRI 2002-03-01T12:17:00
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

Contents of `sessions.json`. Note how the `session_id` `Levels` are clearly described.

```json
{
"participant_id": {
"Description": "BIDS participant identifier"
},
"session_id": {
"Description": "BIDS session identifier",
"Levels": {
"ses-baseline": "Baseline visit for MRI and assessments",
"ses-followupMRI": "6-months after baseline MRI follow-up",
"ses-interview": "1-year after baseline in-person follow-up"
}
},
"acq_time": {
"Description": "When the data acquisition started"
}
}
```

Contents of `participants.tsv`.

```Text
participant_id sex
sub-01 M
sub-02 F
sub-03 F
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

Contents of `phenotype/demographics.tsv`. Measures or features that can change
from session to session belong here especially.

```Text
participant_id session_id age gender race household_income
sub-01 ses-baseline 10 3 4 5
sub-01 ses-followupMRI 10 3 4 5
sub-01 ses-interview 11 4 4 6
sub-02 ses-baseline 9 1 3 3
sub-02 ses-interview 10 1 7 3
sub-03 ses-baseline 11 2 10 4
sub-03 ses-followupMRI 12 5 10 4
Comment thread
ericearl marked this conversation as resolved.
Outdated
```

For more complete examples, see the `pheno00*`
[bids-examples on GitHub](https://github.com/bids-standard/bids-examples/).

[^1]: A session is any logical grouping of imaging and behavioral data consistent
across participants. Session can (but doesn't have to) be synonymous to a visit
in a longitudinal study. In situations where different data types are obtained over
several visits (for example fMRI on one day followed by DWI the day after)
those can still be grouped in one session. Refer to the
[definition of session](../glossary.md#session-entities) for more details.

[^2]: Datetime format and the anonymization procedure are
described in [Units](../common-principles.md#units).
8 changes: 7 additions & 1 deletion src/common-principles.md
Original file line number Diff line number Diff line change
Expand Up @@ -470,7 +470,7 @@ NIfTI header.

### Tabular files

Tabular data MUST be saved as plain-text, tab-delimited values (TSV) files
Tabular data MUST be saved as plain-text, tab-separated values (TSV) files
(with [extension `.tsv`](glossary.md#tsv-extensions)),
that is, [CSV files](https://en.wikipedia.org/wiki/Comma-separated_values) where commas are replaced by tab characters.
Tabs MUST be true tab characters and MUST NOT be a series of space characters.
Expand Down Expand Up @@ -532,6 +532,12 @@ Note that if a field name included in the data dictionary matches a column name
then that field MUST contain a description of the corresponding column,
using an object containing the following fields:

!!! success "Guideline 1"

For [best tabular phenotypic data](./appendices/phenotype.md):
Each tabular phenotypic data TSV file MUST be accompanied by
a corresponding data dictionary JSON file.

<!-- This block generates a metadata table.
The definitions of these fields can be found in
src/schema/objects/metadata.yaml
Expand Down
Loading