Writing Research data
Could you prove what happened to sample A7?
Most labs can tell you what a result was. Far fewer can prove how it got there. The difference is a chain of custody, and it is mostly a recordkeeping problem, not a science problem.
Pick a figure from your last paper. Now trace one data point in it backwards: which sample, which aliquot, which run, which instrument, which method version, which wells were excluded and on whose call, and which analysis produced the number that ended up on the axis.
If you can do that in under a minute, stop reading, because you have solved something most facilities have not. If it would take you an afternoon of hunting through a shared drive, an inherited spreadsheet and someone’s memory, you are in the normal case.
The gap is not rigour, it is recordkeeping
The uncomfortable part is that this has almost nothing to do with how carefully the work was done. Rigorous labs lose provenance constantly, because provenance lives in four places that do not talk to each other:
- Sample identity lives in a freezer sheet, a notebook, or a naming convention only one person fully understands.
- Instrument output lands wherever the instrument’s software puts it, in whatever the vendor’s export looks like this year.
- Analysis happens in a workbook that gets copied, edited and renamed, so the formula that produced a number is not necessarily the formula in the file today.
- Decisions (the exclusions, the re-runs, the threshold that got nudged) happen in conversation and are never written down at all.
Each of those is defensible on its own. Together they mean the record of what happened is reconstructed after the fact rather than captured as it happens. Reconstruction is fine until someone asks a question you cannot answer from memory.
Three moments when it stops being academic
A reviewer asks. Journals increasingly ask for the underlying data and the analysis behind a figure. “Here is the workbook” invites the follow-up question of which version, and what changed between it and the one used at submission.
A funder asks. Data-management and sharing commitments are made at application time and become someone’s job after the award. That someone is usually the facility.
Someone leaves. The postdoc who knew the naming convention takes it with them. This is the most common failure by a wide margin, and the least dramatic: the data is all still there, and no one can use it.
What a chain of custody actually requires
The concept is borrowed from forensics, and the requirement is narrower than it sounds. You need each event that touched a sample to be recorded when it happens, with enough context to be meaningful, in a form nobody can quietly edit later.
In practice that means five things:
- A durable identity per sample and aliquot, including derivations, so A7-2 knows it came from A7.
- Events, not states. Store what happened (“well B4 excluded, QC flag, by whom, at what time”), not just the current value. States overwrite; events accumulate.
- Versioned methods. The analysis that produced a number must name the method version and the parameters, because both drift.
- Tamper-evidence. Each event carries a hash of the previous one, so altering history breaks the chain visibly. This is what makes a record checkable rather than merely trustworthy.
- Open export. If the record cannot leave the system, it is not a scientific record. It is a vendor’s asset.
Notice how little of that is assay-specific. Only the analysis step differs from one experiment type to the next. Identity, events, versioning and tamper-evidence are identical for every experiment your lab runs, which is precisely why solving it per-assay never works.
The honest objection
“We do not have time to log all of that.” Correct, and if capturing provenance is a separate step someone has to remember, it will not happen. It has to be a side effect of doing the work: registering the sample, ingesting the export, running the analysis. The moment it becomes a second system to keep in sync, you have added labour and gained nothing.
That is the actual design constraint. Not “how do we store more metadata”, but “how do we capture the record without adding a single step to the person doing the experiment”.
A test you can run this week
Pick one figure. Ask the person who generated it to reconstruct the full path from freezer to axis, and time it. Then ask what would happen if they were on leave.
Whatever the answer is, it is your current chain of custody. You may be pleasantly surprised. Most people are not, and knowing the number is worth more than any tool you could buy on the strength of a demo.
We are building Lattice around exactly this problem, with ten founding core facilities. If you want the reproduce-your-own-workbook version of this argument rather than the essay, send us one export.