IVF Reads / Are Fertility Apps Reliable?

Are Fertility Apps Reliable?

M
Written by MayaPublished Updated
IVFPulse blog image,
AI summary

Calendar-based fertility apps are unreliable for pinpointing the day of ovulation. In a study where 949 participants collected urine daily for a full cycle so that the LH surge could set the true ovulation day, the accuracy of ovulation prediction across those 949 participants was no better than 21% for the apps tested (Johnson, Marriott and Zinaman, Current Medical Research and Opinion 2018;34(9):1587-94). In a separate assessment of 20 websites and 33 apps against a standardised 28-day cycle, one website and three apps predicted the precise fertile window (Setton, Tierney and Tsai, Obstetrics and Gynecology 2016;128(1):58-63). Apps that use signs from the current cycle rather than the length of past ones score better.

  • Johnson, Marriott and Zinaman (Curr Med Res Opin 2018;34(9):1587-94): of 949 volunteers who collected urine for a whole cycle, 34% believed they had a 28-day cycle but only 15% did; the most likely day of ovulation in a 28-day cycle was day 16, and even that occurred in only 21% of such cycles. App accuracy for the day of ovulation was no better than 21%.
  • Setton, Tierney and Tsai (Obstet Gynecol 2016;128(1):58-63): of 20 websites and 33 apps tested against a standardised 28-day cycle with ovulation on day 15 and a fertile window of days 10-15, one website and three apps predicted the precise fertile window.
  • Freis et al (Front Public Health 2018;6:98): all six calendar-based apps assessed scored zero out of 30; symptothermal apps using signs from the current cycle scored 19 to 20 out of 30. The authors warn that a prediction wrong by a few days may be worse than untimed intercourse.
  • Wilcox, Weinberg and Baird (N Engl J Med 1995;333(23):1517-21), 221 women and 625 cycles: conception occurred only when intercourse took place in a six-day period ending on the estimated day of ovulation, with the probability rising from 0.10 five days before ovulation to 0.33 on the day itself.
  • Bull et al (npj Digital Medicine 2019;2:83), 612,613 ovulatory cycles from 124,648 app users: mean cycle length 29.3 days, mean follicular phase 16.9 days (95% CI 10-30) and mean luteal phase 12.4 days (95% CI 7-17) — the variable part of the cycle is the part calendar maths treats as fixed.
  • ASRM and SREI Practice Committees (Fertil Steril 2022;117(1):53-63): the fertile window is the six-day interval ending on the day of ovulation, intercourse every 1 to 2 days during it yields the highest pregnancy rates, and results with intercourse two to three times weekly are nearly equivalent.
  • Thigpen, Patel and Zhang (J Med Internet Res 2025;27:e60667), 1,155 ovulatory cycles from 964 participants benchmarked against ovulation prediction kits: a sensor-based method detected 1,113 of 1,155 ovulations with a mean error of 1.26 days, against 3.44 days for the calendar method, which performed significantly worse in participants with irregular cycles.

Can a fertility app predict your ovulation day accurately?

Not to the day, if the app works from cycle length. In a study where 949 volunteers collected urine samples across a full menstrual cycle so that the luteinising hormone surge could establish when ovulation actually occurred, the accuracy of ovulation prediction across those 949 participants was no better than 21% for the cycle-tracking apps tested (Johnson, Marriott and Zinaman, Current Medical Research and Opinion 2018;34(9):1587-94). A separate assessment compared 20 websites and 33 apps against a standardised 28-day cycle with ovulation on day 15 and a fertile window of cycle days 10 to 15: one website and three apps predicted the precise fertile window (Setton, Tierney and Tsai, Obstetrics and Gynecology 2016;128(1):58-63).

That is a statement about calendar prediction, not about every app. Tools that read signs from the cycle you are currently in — temperature, cervical fluid, a hormone test — perform differently, and are treated separately below.

Best accuracy achieved by a cycle app for the day of ovulation21%Measured against the urinary LH surge in 949 participants who collected daily urine samples for one complete menstrual cycle. The standard days and rhythm methods were most likely to predict ovulation in the same cohort (70% and 89% respectively) but had very low accuracy.Johnson S, Marriott L, Zinaman M. Curr Med Res Opin 2018;34(9):1587-1594. NCT01577147. The authors are affiliated with SPD Development Company, which makes ovulation tests.

How wide is the target an app is aiming at?

Six days. In 221 women contributing 625 cycles in which the day of ovulation could be estimated from urinary hormones, conception occurred only when intercourse took place during a six-day period ending on the estimated day of ovulation; the probability of conception rose from 0.10 when intercourse happened five days before ovulation to 0.33 on the day of ovulation itself (Wilcox, Weinberg and Baird, New England Journal of Medicine 1995;333(23):1517-21). A later re-analysis that corrected for error in identifying ovulation found the same six-day interval in two independent datasets, with the highest probability on the day before ovulation and a probability close to zero after ovulation (Dunson et al, Human Reproduction 1999;14(7):1835-9).

The ASRM and SREI Practice Committees define the fertile window the same way — the six-day interval ending on the day of ovulation, with peak fecundability in the two days before ovulation (Fertility and Sterility 2022;117(1):53-63). Two things follow. The window opens before ovulation and shuts on it, so an app that flags fertile days after ovulation is flagging days with close to zero probability. And because the window is six days wide, being wrong by three days can place it entirely outside the real one.

Why calendar maths misses

Because it assumes the stable part of the cycle is the part that moves. Among 612,613 ovulatory cycles logged by 124,648 users of one app, mean cycle length was 29.3 days, the mean follicular phase was 16.9 days with a 95% confidence interval of 10 to 30 days, and the mean luteal phase was 12.4 days with an interval of 7 to 17 days (Bull et al, npj Digital Medicine 2019;2:83). The follicular phase — the stretch before ovulation that calendar maths fixes at 14 days — is where almost all the variation lives.

The Johnson data show the same thing from the user's side. Of 949 volunteers, 34% believed they had a 28-day cycle and only 15% actually did. Even within genuine 28-day cycles the most likely day of ovulation was day 16, and that day accounted for only 21% of those cycles among the 949 participants. There is no single day to predict from cycle length, which is why an app that only knows your period dates cannot supply one.

The warning attached to this is not academic. Freis and colleagues note that a fertile window wrong by a few days may lead a couple to concentrate intercourse on less fertile or non-fertile days, which is worse than not timing it at all (Frontiers in Public Health 2018;6:98).

Which app designs scored better, and on what?

Freis and colleagues built a 30-point scoring system covering which parameters an app uses to set the fertile window, what evidence exists for that method, what evidence exists for the app itself, and whether qualified counselling is offered. They then scored 12 apps available in German and English:

  • All six calendar-based apps scored 0 of 30. They set the fertile days from previous cycles and met none of the criteria.
  • The two calculothermal apps scored 3 of 30 and 2 of 30.
  • The four symptothermal apps, which read the current cycle through temperature and cervical fluid, scored 20, 20, 19 and 11 of 30.

The authors judged three of the twelve eligible for further study and called for prospective evaluation by independent investigators free of commercial interest. That caveat is worth carrying into every figure on this page: several of the largest datasets in this field were produced by the companies selling the products, and are marked as such in the source list below.

What are these apps checked against?

Three references, in ascending order of rigour, and knowing which one a claim rests on tells you how much it is worth:

  • Urinary LH. Cheap and widely used, and the benchmark in the Johnson study. ASRM notes that urinary LH devices carry approximately 7% false-positive results, versus a surge that may be followed by ovulation at any point in the next 2 days.
  • Basal body temperature. Confirms ovulation after the fact through the post-ovulatory rise, which makes it useful for learning a pattern and useless for acting on the current day.
  • Serial ultrasound. The reference standard for the actual day of follicle rupture, and the comparator a protocol published in Medicina 2023;59(9):1513 sets out for validating at-home urine hormone monitoring in regular cycles, PCOS and athletes. That paper is a study protocol, so it reports a design rather than results.

An app validated against its own algorithm has been validated against nothing. The question to ask of any accuracy claim is what it was measured against, and in how many cycles.

Not sure whether the timing is the problem?

IVY can read your cycle tracking and test results together and set out what they support, what they do not, and which test would actually answer the question.

Do sensor-based apps do better than calendar apps?

On the evidence retrieved, yes, though the margin is measured in days rather than orders of magnitude. In 1,155 ovulatory cycles from 964 participants, with ovulation prediction kits as the benchmark, a wearable-physiology method detected 1,113 of 1,155 ovulations with a mean error of 1.26 days, against 3.44 days for the calendar method applied to the same cycles (Thigpen, Patel and Zhang, Journal of Medical Internet Research 2025;27:e60667). Accuracy fell for abnormally long cycles, to a mean absolute error of 1.7 days from 1.18.

The finding that matters most in that study is about who the calendar method fails: it performed significantly worse in participants with irregular cycles, while the physiology method did not differ by cycle variability. Participants were recruited from the manufacturer's own commercial database, which is a limitation the authors state.

What happens with irregular cycles and PCOS?

This is where calendar prediction stops being merely imprecise. Three separate failure modes are documented:

  • No surge to find. In the Johnson cohort, no LH surge was detected at all for 99 of the 949 women during the cycle studied. An app extrapolating from period dates will still produce a confident fertile window for such a cycle.
  • Normal-looking length, abnormal luteal phase. In women with polycystic ovaries who reported regular cycles, median cycle length was 28 days and cycle-length variation was no greater than in fertile women with normal ovaries — but early luteal phase progesterone was significantly lower (Joseph-Horne et al, Human Reproduction 2002;17(6):1459-63). A cycle length an app is happy with does not confirm that ovulation went well.
  • Screening features that have not been tested on people. The irregular-cycle feature built into one large tracking app to flag PCOS risk was piloted on 9 virtual test subjects, over-predicted risk compared with a physician, and was reported explicitly as work done before empirical testing on human subjects (Rodriguez et al, JMIR Formative Research 2020;4(5):e15094).

Persistently irregular cycles are a reason to be assessed rather than to track harder. How PCOS affects conception covers what the diagnosis does and does not change, and when ovulation tests read positive without ovulation covers a failure mode no app can see.

Does timing intercourse by app help you conceive?

The honest answer is that regular intercourse removes most of the need for precision. The ASRM and SREI Practice Committees state that intercourse every 1 to 2 days during the fertile window yields the highest pregnancy rates, and that results with less frequent intercourse — two to three times weekly — are nearly equivalent; they advise couples not to restrict frequency when attempting conception (Fertility and Sterility 2022;117(1):53-63).

Read alongside a six-day window, that is the practical conclusion. Intercourse every two to three days across the cycle cannot miss a six-day window, and needs no prediction at all. An app earns its place if it helps you notice a pattern, prompts a hormone test at roughly the right time, or surfaces cycles that look nothing like each other — not by naming a day.

The underlying biology of that window is set out in what the fertile window is and how to track it, and the hormone results an app is indirectly guessing at in what FSH, LH and estradiol actually measure.

What the evidence does not establish

Several claims made for and against these apps are not supported by the sources here.

  • That using a fertility app raises the chance of conceiving. No retrieved trial randomised couples to an app against no app and measured pregnancy. The accuracy studies measure prediction, not outcome.
  • That any app is accurate enough to avoid pregnancy on the same footing as contraception. The one large contraceptive dataset retrieved covers a temperature-based app used for that purpose and reports a typical-use Pearl Index of 6.9 pregnancies per 100 woman-years against 1.0 for perfect use, in a cohort of 22,785 users — with 12-month discontinuation of 54%. That is a different product used for a different purpose, and it was authored by the company that sells it.
  • That the 21% figure applies to every app. It was measured against the cycles of those 949 participants, on the cycle-tracking apps available at the time, none of which published a methodology, which is itself part of the finding.
  • That sensor-based apps are accurate in PCOS. The wearable study reported no loss of accuracy with cycle variability, but did not enrol a diagnosed PCOS group, and the ultrasound-validated PCOS comparison exists so far only as a protocol.
  • That app-recorded data are private. No source retrieved for this article assessed the data-handling practices of these products, so nothing here supports a claim in either direction.

What is established is narrow and usable: the fertile window is six days and ends on the day of ovulation, cycle length alone cannot locate it, apps reading current-cycle signs outperform apps reading past cycle lengths, and intercourse every two to three days makes the prediction question largely moot.

Keep reading

10 Sources

  1. Johnson S, Marriott L, Zinaman M. Can apps and calendar methods predict ovulation with accuracy? Curr Med Res Opin 2018;34(9):1587-1594. 949 volunteers collected urine for one complete cycle; LH measurement assigned the surge day. Mean cycle length 28 days (range 23-35); 34% of women believed they had a 28-day cycle but only 15% did; no LH surge was seen for 99 women; the most likely day of ovulation in a 28-day cycle was day 16, occurring in 21% of such cycles. Accuracy of ovulation prediction was no better than 21% by the apps; the standard days and rhythm methods were most likely to predict ovulation (70% and 89% respectively) but had very low accuracy. NCT01577147. Two authors are affiliated with SPD Development Company, a manufacturer of ovulation tests. Current Medical Research and Opinion
  2. Setton R, Tierney C, Tsai T. The Accuracy of Web Sites and Cellular Phone Applications in Predicting the Fertile Window. Obstet Gynecol 2016;128(1):58-63. Twenty websites and 33 apps assessed against a standardised 28-day cycle with estimated ovulation on cycle day 15 and a fertile window of cycle days 10-15. One website and three apps predicted the precise fertile window; all included the most fertile cycle day, but the range of the predicted window varied widely. The authors state the clinical effect of this inaccuracy is unknown. Obstetrics and Gynecology
  3. Freis A, Freundl-Schutt T, Wallwiener LM, Baur S, Strowitzki T, Freundl G, Frank-Herrmann P. Plausibility of Menstrual Cycle Apps Claiming to Support Conception. Front Public Health 2018;6:98. Twelve apps scored against a 30-point system. All six calendar-based apps scored zero; the two calculothermal apps scored 3 and 2 of 30; the four symptothermal apps scored 20, 20, 19 and 11 of 30. The authors state that a deviation of a few days may lead a couple to focus on less fertile or non-fertile days and may therefore be worse than random intercourse, and call for prospective studies by investigators free of commercial bias. Frontiers in Public Health
  4. Practice Committee of the American Society for Reproductive Medicine and the Practice Committee of the Society for Reproductive Endocrinology and Infertility. Optimizing natural fertility: a committee opinion. Fertil Steril 2022;117(1):53-63. Defines the fertile window as the six-day interval ending on the day of ovulation, with peak fecundability in the two days before ovulation; states that intercourse every 1 to 2 days during the fertile window yields the highest pregnancy rates and that two to three times weekly is nearly equivalent; advises couples not to restrict frequency; notes that urinary LH detection devices carry approximately 7% false-positive results and that ovulation may occur at any time within the 2 days after the surge; and cites a maximum accuracy of 21% for downloadable calendar apps in predicting the day of ovulation. American Society for Reproductive Medicine
  5. Wilcox AJ, Weinberg CR, Baird DD. Timing of sexual intercourse in relation to ovulation. N Engl J Med 1995;333(23):1517-1521. 221 healthy women planning pregnancy collected daily urine specimens; in 625 cycles with an estimable ovulation date, 192 pregnancies were initiated. Conception occurred only when intercourse took place during a six-day period ending on the estimated day of ovulation, with the probability of conception ranging from 0.10 five days before ovulation to 0.33 on the day of ovulation. New England Journal of Medicine
  6. Dunson DB, Baird DD, Wilcox AJ, Weinberg CR. Day-specific probabilities of clinical pregnancy based on two studies with imperfect measures of ovulation. Hum Reprod 1999;14(7):1835-1839. Re-analysis of a 1950s-60s London natural family planning cohort (ovulation identified by basal body temperature shift) and a 1980s North Carolina cohort (urinary hormone assays). After correcting for error in identifying ovulation, the same six-day fertile interval was estimated in both, with the highest probability of pregnancy on the day before ovulation and probabilities close to zero after ovulation. Human Reproduction
  7. Bull JR, Rowland SP, Scherwitzl EB, Scherwitzl R, Danielsson KG, Harper J. Real-world menstrual cycle characteristics of more than 600,000 menstrual cycles. npj Digit Med 2019;2:83. 612,613 ovulatory cycles from 124,648 users of the Natural Cycles app. Mean cycle length 29.3 days; mean follicular phase 16.9 days (95% CI 10-30); mean luteal phase 12.4 days (95% CI 7-17). The authors conclude that identifying the fertile period requires tracking physiological parameters such as basal body temperature and not just cycle length. Funded by Natural Cycles Nordic AB; two authors are employees and two are founders. npj Digital Medicine
  8. Thigpen N, Patel S, Zhang X. Oura Ring as a Tool for Ovulation Detection: Validation Analysis. J Med Internet Res 2025;27:e60667. 1,155 ovulatory cycles from 964 participants recruited from the manufacturer's commercial database, with ovulation prediction kits as the benchmark. The physiology method detected 1,113 of 1,155 ovulations (96.4%) with an average error of 1.26 days, against 3.44 days for the calendar method (P<.001). Accuracy was lower for abnormally long cycles (mean absolute error 1.7 days versus 1.18 days). The calendar method performed significantly worse in participants with irregular cycles; the physiology method did not differ by cycle variability. Journal of Medical Internet Research
  9. Joseph-Horne R, Mason H, Batty S, White D, Hillier S, Urquhart M, Franks S. Luteal phase progesterone excretion in ovulatory women with polycystic ovaries. Hum Reprod 2002;17(6):1459-1463. Urinary pregnanediol-3-glucuronide measured from day 10 to the next menses across three consecutive cycles. Median cycle length was 28 days (range 23-47) in the polycystic-ovary group. Women with polycystic ovaries did not have more variation in cycle length than fertile women with normal ovaries, but had significantly lower progesterone in the early luteal phase. Human Reproduction
  10. Rodriguez EM, Thomas D, Druet A, Vlajic-Wheeler M, Lane KJ, Mahalingaiah S. Identifying Women at Risk for Polycystic Ovary Syndrome Using a Mobile Health App: Virtual Tool Functionality Assessment. JMIR Form Res 2020;4(5):e15094. Pilot of an irregular-cycle feature generating a PCOS risk score in the Clue app, run on 9 virtual test subjects with a physician as the gold standard, explicitly conducted before empirical testing on human subjects. The feature over-predicted PCOS risk relative to the physician; the correlation with physician scores was 0.82 (P=.01) in the first iteration and 0.73 (P=.03) in the second. Three authors were employees of Clue at the time. JMIR Formative Research

Frequently asked questions

Common questions on this topic.

Which is more accurate, an app or an ovulation predictor kit?

The kit, because it measures something happening now. ASRM and SREI state that urinary LH devices detect the surge with approximately 7% false-positive results and that a randomised controlled trial found these devices decrease time to pregnancy; no equivalent trial result was retrieved for a calendar app.

Can an app tell me whether I ovulated at all?

Not on its own. A temperature rise recorded after the fact is consistent with ovulation having occurred, but in the Johnson cohort no LH surge was detected for 99 of 949 women in the cycle studied, and an app working from dates alone will still display a fertile window for a cycle like that.

Is basal body temperature worth recording if it only confirms ovulation afterwards?

It is worth recording to learn a pattern across several cycles rather than to act on today. The authors of the 612,613-cycle dataset concluded that identifying the fertile period requires tracking physiological parameters such as basal body temperature and not just cycle length.

Why does my app move my predicted ovulation day after I log a period?

Because a calendar algorithm recalculates from your recorded cycle lengths. Since the follicular phase carries almost all of the variation — mean 16.9 days with a 95% confidence interval of 10 to 30 days across 612,613 cycles — a retrospective adjustment of the predicted day is the algorithm working as designed, not a fault.

Does a longer or shorter cycle change where the fertile window sits?

Yes, and unpredictably. Mean cycle length in the Johnson cohort was 28 days with a range of 23 to 35, and the day of ovulation varies considerably for any given cycle length, which is why the authors concluded that methods using cycle length alone cannot accurately predict it.

Should I stop using a tracking app?

Nothing retrieved here supports telling anyone to stop. The reasonable use is as a record rather than a forecast, and ASRM and SREI advise against restricting intercourse frequency in order to save it for predicted days.

How long should I use an app before seeing a doctor?

That question is answered by how long you have been trying and by age, not by the app. Persistently irregular cycles, or no cycle at all, are a reason for assessment regardless of how long you have been tracking.