Writing Research data
What a data-management plan actually asks of a core facility
A DMS plan is written by a PI at application time and becomes someone else’s job after the award. Usually the core’s. Here is a practical read on what that means in daily operation.
Data-management and sharing plans are written at application time, by someone optimising for the science and the page limit. They are then made true, or not, by the people who handle the samples, run the instruments and hold the exports. In multi-lab environments, that is very often the core facility, which was not in the room when the commitments were made.
Two things follow. First, the burden is real but narrower than the anxiety around it suggests. Second, almost all of it is recordkeeping you would want anyway.
Check the current requirements and effective dates directly at the source before you plan around them. NIH has been simplifying the DMS plan format (see notice NOT-OD-26-046), and policy detail changes faster than any summary written by a vendor, including this one.
The four commitments that land on the facility
Whatever the format, plans make promises in roughly four areas, and these are the ones with operational consequences downstream:
1. What data will be preserved and shared. Someone must be able to identify, later, which files constitute “the data” for a given study, as distinct from the several hundred other files the instrument produced that week.
2. Metadata and documentation standards. The data must be intelligible to someone who was not there. In practice this means the sample it came from, the method and parameters used, and the instrument and software versions.
3. Retention and access timelines. Data must survive for a stated period and be findable in that window, which is usually longer than the tenure of the person who generated it.
4. Repository or archive destination. Where the data ends up, in a format the repository accepts.
None of that requires a new instrument or a change in how the science is done. All of it requires the record to exist at the moment the work happens, rather than being reconstructed years later.
Where facilities actually get caught
In conversations with directors, the same three failure modes come up:
- Findability, not storage. The data has been kept. It is in a folder structure that made sense to one person in 2024, with filenames that encode a convention nobody wrote down. Kept and unfindable is functionally the same as lost when someone asks for it.
- Missing linkage. The export exists and the sample list exists, but nothing binds a specific file to a specific specimen and a specific method version. That link is the single most valuable piece of metadata, and it is the one most often held only in a person’s head.
- Silent method drift. The analysis parameters used in month 2 are not the ones used in month 14, and no record says when the change happened or why. This is the part that makes a dataset unreproducible even when every byte was preserved faithfully.
A checklist that takes an afternoon
For a facility that wants to be defensible without buying anything:
- Name a data owner per service. One person accountable for what “the data” means for that instrument, even if the work is shared.
- Fix one identifier scheme for samples and aliquots, and use it in filenames, sheets and exports. Consistency beats elegance.
- Write the method down and version it: parameters, thresholds, exclusion policy. A one-page document with a version number is enough.
- Record exclusions and re-runs where the data lives. A decision that only exists in a Slack thread did not happen, as far as the record is concerned.
- Prefer open formats at the point of export wherever an open standard exists for your instrument. Repositories want them, and vendor formats age badly.
- Test a retrieval. Pick a study from a year ago and try to assemble everything the plan promised to preserve. Time it. That number is your real compliance posture.
Step six is the only one that tells you the truth, and it is the one nobody runs.
What software can and cannot do here
Be sceptical of any tool claiming to handle your DMS plan. Software does not make commitments to a funder; your institution does. What a system can do is make the commitments cheap to keep: capture the sample-to-result linkage as a side effect of the work, version methods automatically, keep an audit trail of exclusions and re-analyses, and export the whole thing in open formats when the repository asks.
That is the honest scope. Anything more is a claim about someone else’s obligations, which is not a claim we are willing to make.
Building this into a system of record for labs: Lattice. Also relevant: Could you prove what happened to sample A7?