Source / Clinical scoring paper

Apgar's newborn assessment paper

Published in 1953 from Columbia-Presbyterian's Sloane Hospital for Women in New York, Virginia Apgar's “A Proposal for a New Method of Evaluation of the Newborn Infant” proposed a zero-to-ten score recorded one minute after birth.

The paper mattered because it fixed what delivery-room staff should observe and when, allowing newborn condition to be recorded and compared. Its one-minute research method, however, was not identical to the later one- and five-minute clinical routine, and a score was never a complete diagnosis or an instruction to delay resuscitation.

Clinical setting, 1949–1952

The score answered a delivery-room measurement problem.

Apgar was an anaesthesiologist observing hospital births at a time when maternal sedatives and general anaesthesia could suppress a newborn's breathing. Resuscitation methods were debated, but evidence was difficult to compare because observers used inconsistent endpoints and descriptive labels.

Earlier labels were unstable

The paper criticised “breathing time” and “crying time”: a medicated infant might gasp, stop breathing, and resume later, while an injured infant might never establish what an observer considered a satisfactory cry. “Mild,” “moderate,” or “severe depression” also left wide room for judgment. Here depression meant reduced vital activity, not a psychiatric condition.

Fixed timing made the observations portable

Between 1949 and 1952, Apgar tested observable signs and settled on five that could be judged quickly without special equipment or interrupting care. The original paper specified sixty seconds after the baby's complete birth. It initially relied on two observers where possible, then accepted ratings made by the anaesthesia resident attending the delivery.

The newborn became a measure of obstetric practice

The score was designed not only to classify an infant but to compare maternal pain relief, anaesthetic techniques, modes of delivery, and resuscitation. That purpose places the article within post-war obstetric anaesthesia and a wider, international network of paediatricians, physiologists, anaesthetists, obstetricians, nurses, and midwives studying the transition to breathing.

Reading the 1953 evidence

A simple table rested on a partial hospital series.

During the paper's seven-and-a-half-month reporting period, 2,096 infants were born at Sloane. Scores were available on 84 per cent of the anaesthesia records. Apgar acknowledged that many missing records concerned births with pudendal block or “natural childbirth”—the cases she regarded as the best control group.

Five signs became one number

Heart rate, breathing, response to stimulation, tone, and colour each received zero, one, or two points. Apgar treated heart rate as the most important sign and colour as “by far the most unsatisfactory”; it generated the most disagreement among observers. The total compressed several different physiological observations, which made recording fast but could hide why two infants received the same score.

Group comparisons were exploratory

The paper tabulated scores by mode of delivery and anaesthesia. Among 141 scored caesarean births, for example, 83 infants whose mothers received spinal anaesthesia averaged 8.0, while 54 after general anaesthesia averaged 5.0. These were not randomized groups. Apgar noted maternal and fetal differences that the comparison did not capture, and elsewhere called subgroup results too small or statistically non-significant.

Survival supported classification, not certainty

In the outcome table, nine of 65 infants scoring 0–2 died, compared with two of 182 scoring 3–7 and one of 774 scoring 8–10. Those figures established an association within the reported series; they did not show that the score caused better survival or could predict an individual child's future. The article also excluded infants under 500 grams from its prematurity analysis under the period category “abortions,” a classification that should not be silently mapped onto present-day viability thresholds.

Publication and afterlife, 1952–1966

The familiar routine was assembled after the first paper.

The article followed Apgar's September 1952 presentation to the Twenty-Seventh Annual Congress of Anesthetists. Publication supplied a table, a common observation time, and grouped results that other hospitals could reproduce. The paper was advocacy as well as evidence: its author wanted more precise study of newborn resuscitation and closer attention during the first minute.

A 1958 second report by Apgar, Duncan A. Holaday, L. Stanley James, Irvin M. Weisbrot, and nurse Cornelia Berrien analysed scores from 15,348 infants and related them to mortality, obstetric and anaesthetic circumstances, and umbilical-blood findings. In the 1960s, five-minute scoring became part of the system because it correlated more closely with survival. This expanded evidence and multi-department collaboration—not the 1953 table alone—helped turn a local research tool into a clinical and statistical standard.

The name also acquired a teaching device after publication. In 1962 paediatricians Joseph Butterfield and Mervyn Covey proposed “Appearance, Pulse, Grimace, Activity, Respiration” as a mnemonic. It is a retrospective acronym made from Apgar's surname, not the origin of her five categories. By 1966, Apgar's own reflections and advice warned that scores varied across observers and institutions and recommended a trained scorer other than the person conducting the delivery.

Meaning and limits

A score records condition at a moment; it does not explain every cause.

The article's word asphyxia belonged to a broad contemporary language for newborns with impaired breathing and oxygenation. Current professional guidance is narrower: an Apgar score alone neither diagnoses an intrapartum hypoxic–ischaemic event nor proves that “asphyxia” caused later disability. Prematurity, maternal medication or anaesthesia, congenital conditions, trauma, infection, resuscitation, and ordinary physiological transition can all influence the number.

Nor should care wait for the first score. Apgar explicitly corrected that misunderstanding in 1966, saying resuscitation should begin when needed before the sixty-second observation. Present guidance likewise treats the score as a report of the newborn's status and response to care, not as the trigger that determines the initial steps of resuscitation.

Three components—colour, tone, and reflex response—require interpretation. The original “completely pink” criterion centred light skin as the visual norm, while Apgar herself regarded colour as the weakest sign. Recent historical analysis argues that the problem is broader than wording alone: population-level associations have repeatedly been overextended into diagnoses and predictions about individual infants, with the potential to compound racial inequities.

The primary paper is valuable precisely because it preserves its setting and limits. It records clinician categories and aggregated outcomes, not the experiences of mothers or newborns, and it does not describe consent or long-term follow-up. Apgar thanked nurse Rita Ruane for technical assistance, but the article gives little account of the routine nursing and record work on which scoring depended. It therefore documents an influential proposal, not a neutral or complete history of how newborn assessment changed worldwide.

Across the collection

Continue from the Apgar paper

Virginia Apgar

Follow her work in anaesthesiology, obstetric care, teaching, research, and later public-health advocacy.

The Apgar score

Place the 1952 presentation, 1953 publication, and later clinical adoption on the chronological spine.

Obstetrics and midwifery

Connect newborn assessment to hospital birth, maternal treatment, birth attendants, and changing measures of safety.

References

Primary sources and further reading

  1. Virginia Apgar, “A Proposal for a New Method of Evaluation of the Newborn Infant”

    Current Researches in Anesthesia and Analgesia 32, no. 4 (1953): 260–267. Complete digitised article from the U.S. National Library of Medicine; the central contemporary source for the method, hospital series, comparisons, and authorial claims.

  2. Virginia Apgar, Duncan A. Holaday, L. Stanley James, Irvin M. Weisbrot, and Cornelia Berrien, “Evaluation of the Newborn Infant—Second Report”

    JAMA 168, no. 15 (1958): 1985–1988. DOI: 10.1001/jama.1958.03000150027007. Follow-up study of 15,348 infants; the publisher record and abstract identify the larger clinical team, findings, and institutional affiliations.

  3. Joseph Butterfield and Mervyn J. Covey, “Practical Epigram of the Apgar Score”

    JAMA 181, no. 4 (1962): 353. DOI: 10.1001/jama.1962.03050300073025. Contemporary letter introducing the English-language mnemonic based on Apgar's name.

  4. Virginia Apgar, “The Newborn (Apgar) Scoring System: Reflections and Advice”

    Pediatric Clinics of North America 13, no. 3 (1966): 645–650. DOI: 10.1016/S0031-3955(16)31874-0. A retrospective primary source on the method's original aims, timing, observer variation, later use, and misuse.

  5. Rachel McAdams, “Learning to Breathe: The History of Newborn Resuscitation, 1929 to 1970”

    PhD thesis, University of Glasgow, 2009. An archive-based study of British and American resuscitation debates, physiological research, clinical networks, and the limitations of linear progress narratives.

  6. Rebecca L. Jackson, “The Apgar Score and Race: Why Healthy Babies Are Supposed to Be ‘Pink’”

    History and Philosophy of the Life Sciences 47, no. 4 (2025), article 45. DOI: 10.1007/s40656-025-00693-3. Open-access historical analysis of colour, changing uses of the score, group-to-individual inference, and racialized measurement.

  7. American Academy of Pediatrics Committee on Fetus and Newborn and American College of Obstetricians and Gynecologists Committee on Obstetric Practice, “The Apgar Score”

    Pediatrics 136, no. 4 (2015): 819–822; reaffirmed by ACOG in 2025. DOI: 10.1542/peds.2015-2651. Current professional context on scoring, resuscitation, subjectivity, and inappropriate diagnosis or individual prognostication.