Collecting Women's Health Data in Clinical Practice: A Guide for Clinicians, Clinic Owners, and Research Directors
Stephanie Karzon Abrams, Clinical neuropharmacologist | Clinical Director Mystic Health | Founder, Beyond Consulting | Research Director, Microdosing Collective 501c3
Most clinics are sitting on valuable data; data that can inform how care is provided, what medication labels say, and what future clinical trials get funded to study!
Patients tell you important things like her ketamine infusion hit differently the week before her period. Another says her anxiety spikes every month, predictably, then lifts. A third reports that the session she found most difficult happened to fall in the days just before menstruation, and another told you that the her period came a day after her psilocybin she received in her session; a whole 4 days earlier than usual. Perhaps you noted it, perhaps you acknowledged it and moved on.
This guide is for the clinicians, clinic owners, and research directors who are ready to ask important questions and most importantly ensure that it is being collected in a standardized way so it can be used for the greater good of research and informing clinical practice!
The critical requirement is a common data model — standardized data elements, definitions, and collection procedures across all sites. Without this, aggregation produces noise. Professional society registries (STS cardiac surgery databases, ACC's NCDR, AHA's Get With The Guidelines) have succeeded precisely because they enforce granular, consistent data standards across hundreds of participating institutions, often funded by hospital participation fees rather than traditional grants.
Why Cycle Phase Data Is Clinical Data
The relationship between the menstrual cycle and symptom severity is not anecdotal. Estrogen regulates serotonin synthesis, receptor sensitivity, and degradation rate; progesterone converts to allopregnanolone, a neurosteroid that modulates GABA-A receptors with anxiolytic properties; and both hormones interact with the glutamatergic system, which ketamine targets directly. When estrogen drops sharply in the late luteal phase, the serotonergic system changes. When allopregnanolone falls, the anxiolytic effect disappears. These are measurable neurochemical events.
A patient arriving for a psilocybin session in the late luteal phase is in a neurochemically different state than the same patient in the follicular phase. Whether that difference affects their session outcome is a question we do not yet have controlled trial data to answer. What we do have is the infrastructure to start collecting that data, right now, in clinical settings, using validated instruments that already exist.
The Core Problem with Retrospective Recall
Ask a patient whether her symptoms vary with her cycle and she will almost certainly say yes. Ask her to describe the pattern and she will give you an answer shaped by the worst days she remembers, the most recent cycle, and the way memory conflates emotional intensity with frequency. This is not a failure of self-awareness. It is how retrospective recall works.
The Moos MDQ Form C — a widely used retrospective premenstrual symptom measure — overestimates symptom severity relative to prospective daily ratings and is generally excluded from regulatory-quality research for exactly this reason. The American College of Obstetricians and Gynecologists and the DSM-5 both require a minimum of two consecutive months of prospective daily ratings to confirm that symptoms are genuinely cycle-phase-dependent rather than reflecting recall bias or misattribution.
If the data you are collecting is retrospective, it is clinically useful as a starting point and scientifically unreliable as anything more.
Step One: Choose a Validated Daily Rating Instrument
The gold standard for prospective cycle-phase symptom tracking is the Daily Record of Severity of Problems, or DRSP. It is a 24-item daily diary with a 6-point severity scale ranging from 1 (not at all) to 6 (extreme), covering 21 symptom items and 3 functional impairment items. Both ACOG and the DSM-5 use it as a reference instrument for PMDD diagnosis confirmation, and the Carolina Premenstrual Assessment Scoring System, the C-PASS, provides a standardized algorithm for converting DRSP data into DSM-5-compliant diagnoses with 98% agreement with expert clinician diagnosis.
For clinics that need something lighter: the Calendar of Premenstrual Experiences, the COPE, covers 22 items on a 4-point scale and is well-validated. Visual Analogue Scales are appropriate for single-symptom tracking when that is all you need. The Premenstrual Tension Syndrome Rating Scale is another validated option used across the literature.
What these instruments share, and what makes them useful, is daily completion. The patient fills them out each day, not at the end of the month, not before the appointment. Daily. That requirement is non-negotiable if the data is to mean anything, and it is the thing most clinics do not build infrastructure to support.
A digital daily diary embedded in your patient portal, sent via app notification, or even delivered by text message takes this from a clinical aspiration to a clinical workflow. The friction of daily completion is the single largest implementation barrier, and it is solvable.
Step Two: Define Cycle Phase Accurately
This is where most cycle research goes wrong, including research conducted by well-resourced academic teams. A comprehensive methodological review of 146 studies identified six different methods for defining menstrual cycle phase, with poor consistency across them. Assuming a standard 28-day cycle and scheduling visits accordingly leads to substantial phase misclassification because of inter- and intra-individual variability in cycle length. Scheduling visits by set cycle days alone captures the LH surge, the reliable marker of ovulation, only 37 to 57% of the time even with perfect patient attendance.
The most rigorous option is daily urinary hormone sampling measuring estrone-3-glucuronide and LH, or serum estradiol and progesterone, to confirm ovulation and define follicular versus luteal phases biochemically. This is research-grade infrastructure; most clinical practices will not implement it routinely, but research directors designing studies should.
Tools like the Mira or Inito can collect urine samples easily. For clinical practice, a home fertility monitor detecting the LH surge is a good and practical middle ground. Clearblue and similar devices are widely available, inexpensive relative to the clinical value of the data, and give you a reliable marker of the ovulatory window that self-reported cycle dates cannot. The minimum viable approach is self-reported menses onset plus luteal-phase confirmation via serum progesterone above 16 nmol/L. This is workable in a busy clinical setting. It is less precise than LH surge detection and substantially more reliable than calendar date alone. What you should not do is assume. A patient with a 25-day cycle and a patient with a 34-day cycle are at very different points in their hormonal arc on day 20. Treating them as equivalent because their calendar day matches is a misclassification, and that misclassification compounds across a patient population until your data tells you nothing useful about cycle phase at all.
Step Three: Build It Into Your Intake
The clinical application of all of this is simpler than it sounds, and it starts at intake.
Add cycle history questions to your standard intake form. Not just "are you currently menstruating?" but: what is your average cycle length, do your symptoms vary with your cycle, have you ever been evaluated for PMDD or PMS, are you currently using hormonal contraception (which suppresses the natural cycle entirely and renders cycle-phase tracking irrelevant unless you are tracking synthetic hormone fluctuations), and critically, where are you in your cycle today?
That last question costs you nothing and immediately gives you context about the patient sitting across from you.
For patients who will be receiving ketamine-assisted therapy, psilocybin therapy, or any extended therapeutic engagement, consider adding the DRSP as a two-month pre-treatment baseline measure. Two months of daily ratings before the first session gives you a documented symptom profile, a confirmed cycle-phase pattern if one exists, and a baseline against which to measure treatment effects. It also identifies patients who may have undiagnosed PMDD, a clinically meaningful finding that should inform session timing, preparation protocols, and integration support.For patients already in treatment, adding a brief daily check-in that includes cycle day — even just "what day of your cycle is today, estimated" — creates longitudinal data over time that can be analyzed later even if it was not collected with a formal protocol from the start. Imperfect prospective data is better than perfect retrospective recall.
Step Four: What to Do With the Data
Collecting cycle-phase data without a plan for using it is record-keeping theater. Here is what it enables in practice.
Session timing: If a patient's symptom burden is consistently highest in the late luteal phase, there is a clinical argument for scheduling acute psychedelic sessions outside that window. The follicular phase, characterized by rising estrogen, lower anxiety, and greater psychological openness for many patients, may represent a more favorable neurochemical context for a first or particularly challenging session. This is not a rigid protocol. It is a consideration that the data makes possible.
Preparation and anticipatory support: Knowing that a patient will be in the late luteal phase around the time of an upcoming session allows you to adjust preparation accordingly to address what she is likely to bring into the room, not what you assume she will bring based on her general presentation.
Integration timing:The week after a ketamine or psilocybin session tends to be a period of heightened neuroplasticity and emotional openness. If that week overlaps with the luteal phase for a patient with significant premenstrual symptoms, integration sessions need to account for that — the material may be more volatile, the emotional processing more intense, and the window for durable behavioral change more complex to work with than in a patient whose post-session week falls in the follicular phase.
Outcome tracking: If you are collecting symptom data daily over time, you can begin to ask whether your treatments are affecting cycle-phase-dependent symptom severity, not just general measures of depression or anxiety. This is a more precise question. It is also a more clinically meaningful one for the patients you are treating.
A Note for Research Directors
If you are designing a study rather than building a clinical protocol, the methodological standards are higher and the regulatory requirements are specific.
The FDA's 2009 PRO Guidance and the SPIRIT-PRO Extension published in JAMA in 2018 define what an IRB and a regulator need to see from patient-reported outcome data. Content validity and establishing that the instrument measures what it claims to measure through patient focus groups and cognitive interviewing must be documented. The PRO hypothesis must be prespecified in the statistical analysis plan before data collection begins; you cannot decide after the fact that cycle phase was your primary variable of interest. The FDA generally recommends no more than a one-day recall period for PROs, which aligns perfectly with daily cycle diaries. A recent review found that prespecified sensitivity analyses for missing PRO data were documented in fewer than 30% of pivotal trials; an oversight that creates significant regulatory vulnerability.
The minimal clinically important difference, the MCID, must be calculated for each scale to establish what constitutes a meaningful change rather than just a statistically significant one. And for any study examining cycle-phase-dependent effects, realignment methods, statistical techniques that align data to the LH surge rather than to calendar days, followed by multiple imputation for missing data points, are the appropriate analytical approach.
A 2024 analysis of more than 145,000 symptom logs from the MenoLife app demonstrated that app-based daily symptom tracking can identify distinct symptom clusters across reproductive life stages at the scale needed for this kind of research. Digital infrastructure that already exists in patient populations can be leveraged. The gap, as ever, is not the tools. It is the study design.
The Minimum Viable Clinical Protocol
For clinics that want to start now without building a research program, here is the floor.
Add cycle history to intake. Include current cycle day, average cycle length, history of PMDD or significant premenstrual symptoms, and current hormonal contraception status. Use the DRSP as a two-month pre-treatment baseline for patients entering extended therapeutic programs. Add cycle day to your daily or weekly patient check-in. Use home LH surge detection for patients where session timing relative to cycle phase is clinically relevant. Document it all in a way that can be analyzed later.
Those five moves do not require a research budget, a data scientist, or a new EMR. They require intention and a revised intake form.The data you collect this year, even imperfectly, is data you will not have to reconstruct in five years when the field finally catches up to the question.
A Note on Data Ownership
HIPAA governs individually identifiable health information held by covered entities (hospitals, clinicians, insurers). De-identified data can be shared for essentially any purpose without patient consent. However, re-identification risk is increasingly recognized as a real concern given modern data-linking capabilities.
The Common Rule (updated 2018) requires that research consent forms disclose whether identified data may be de-identified and reused for future studies without additional consent. Secondary analysis of existing de-identified healthcare databases is generally exempt from informed consent requirements.
In practice, the entity that controls the data infrastructure typically controls access. If a professional society or nonprofit builds the registry, it sets the data use agreements. If a commercial platform (e.g., TriNetX) hosts the data, the platform's terms govern access. If individual practices collect data using a shared addendum but store it locally, a federated model preserves local data ownership while enabling centralized queries.
The emerging best practice is a Data Governance Board with patient representation — the FDA itself has called for "patient-mediated data-sharing" models where patients voluntarily share EHR data with a coordinating center controlled by a patient-elected board. [9] The Diabetes Action Canada network provides a working example: governance is built on eight principles (transparency, accountability, participation, etc.) with a Research Governing Committee whose majority are patients or data-contributing physicians.
Galilea® and Cycle-Aware Clinical Practice
The Galilea® framework was designed, in part, as a response to the absence of this kind of systematic thinking in integrative and psychedelic medicine. It provides clinicians and clinic owners with the tools, protocols, and training to build hormone-aware care into practice at the level of intake design, session timing, preparation, and integration. This real-world data can also contribute to demystifying some of the unanswered questions in women’s health and create strong signal to fund prospective clinical trials that will inform care in a valuable way.
The concept of a learning health system model applied to sex-hormone-drug interactions is feasible, increasingly accepted by regulators, and potentially faster than traditional trials. But the governance decisions made at the outset like who designs the data model, who hosts the infrastructure, who controls access, and whether the effort is classified as research or QI will determine whether the resulting dataset is a regulatory asset or an interesting but unusable collection of observations.
If you are building this into your practice and want a framework to work from, Galilea® The Method is where to start.
Reach out!
References
[1] Endicott J, et al. Daily Record of Severity of Problems. Arch Women's Ment Health. 2006.
[2] American College of Obstetricians and Gynecologists. Premenstrual Syndrome. ACOG Practice Bulletin No. 15. 2000.
[3] DSM-5. Premenstrual Dysphoric Disorder diagnostic criteria. American Psychiatric Association. 2013.
[4] Yonkers KA, et al. Premenstrual syndrome. Lancet. 2008.
[5] Eisenlohr-Moul TA, et al. C-PASS scoring system. Psychol Assess. 2017.
[6] Freeman EW. Premenstrual syndrome and premenstrual dysphoric disorder. J Womens Health. 2003.
[7] Schmalenberger KM, et al. Menstrual cycle phase classification methods. Psychoneuroendocrinology. 2021.
[8] Toffol E, et al. Serum hormone confirmation in cycle research. Gynecol Endocrinol. 2014.
[9] Stern K, et al. LH surge detection and realignment methods. Horm Behav. 2018.
[10] Bull JR, et al. Real-world menstrual cycle characteristics from a mobile app. NPJ Digit Med. 2019.
[11] Prior JC, et al. Luteal phase progesterone thresholds. Endocr Rev. 1990.
[12] FDA. Guidance for Industry: Patient-Reported Outcome Measures. 2009.
[13] Calvert M, et al. SPIRIT-PRO Extension. JAMA. 2018.
[14] Patrick DL, et al. Content validity methods for PRO instruments. Value Health. 2011.
[15] Gnanasakthy A, et al. Recall period recommendations for PRO instruments. J Patient Rep Outcomes. 2019.
[16] Efficace F, et al. Missing PRO data in clinical trials. Lancet Oncol. 2020.
[17] Bell ML, et al. Sensitivity analyses for missing PRO data. Trials. 2018.
[18] Bruinvels G, et al. MenoLife app symptom log analysis. NPJ Women’s Health. 2024.
[19] Principles for Health Information Collection, Sharing, and Use: A Policy Statement From the American Heart Association. Circulation. 2023. Spector-Bagdady K, Armoundas AA, Arnaout R, et al.Guideline
[20]Bridge the gap: The need for harmonized regulatory and ethical standards for postmarketing observational studies. Pharmacoepidemiology and Drug Safety. 2017. Urushihara H, Parmenter L, Tashiro S, Matsui K, Dreyer N.Review
[21]Data Governance Requirements for Distributed Clinical Research Networks: Triangulating Perspectives of Diverse Stakeholders. Journal of the American Medical Informatics Association : JAMIA. 2013. Kim KK, Browe DK, Logan HC, et al.
[22] Premarket and Postmarket Real-World Evidence Studies Supporting U.S. Food and Drug Administration Regulatory Decision-Making, 2016-2024. Clinical Trials. 2026. Li LY, Ramachandran R, Ross JS, Wallach JD
[23] TriNetX Dataworks‐ USA : Overview of a Multi‐Purpose, De‐Identified, Federated Electronic Health Record Real‐World Data and Analytics Network and Comparison to the US Census. Pharmacoepidemiology and Drug Safety. 2025. Stein E, Hüser M, Amirian ES, Palchuk MB, Brown JS

