BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201805Z
LOCATION:241
DTSTART;TZID=America/Chicago:20251117T111000
DTEND;TZID=America/Chicago:20251117T113000
UID:submissions.supercomputing.org_SC25_sess205_ws_cafcw114@linklings.com
SUMMARY:PathLlama: A Language Model for Automated Cancer Surveillance
DESCRIPTION:Patrycja Krawczuk, John Gounley, Abhishek Shivanna, and Mayank
 a Chandrashekar (Oak Ridge National Laboratory (ORNL)); Elizabeth Hsu (Nat
 ional Cancer Institute); and Heidi Hanson (Oak Ridge National Laboratory (
 ORNL))\n\nTransforming unstructured information into structured common\nda
 ta models (CDM) is a critical step for enabling cancer\nsurveillance and a
 dvancing precision medicine. CDMs standardize\nthe structure and content o
 f oncologic data extracted\nfrom electronic health records. Unfortunately,
  traditional Extract\nTransform Load processes for electronic health data 
 capture\nare generally rule-based, error-prone, and produce static\ndatase
 ts unsuitable for near real-time information retrieval.\n\nThe Modeling Ou
 tcomes using Surveillance Data and Scalable\nAI for Cancer (MOSSAIC) proje
 ct developed and deployed\na hierarchical self-attention (HiSAN) model cap
 able\nof autocoding approximately 30% of National Cancer Institute\nSurvei
 llance, Epidemiology, and End Results (SEER) registry\ncancer pathology re
 ports [1], [2]. While a significant step\nforward, this falls short of the
  broader goal of automatically\ncoding all pathology reports. Fully automa
 ting CDM conversion\nwould facilitate clinical trial matching, decision su
 pport\ndashboards, real-time case ascertainment, and population\nhealth su
 rveillance.\n\nThe distribution of cancer phenotypes in real-world data is
 \nhighly imbalanced. While HiSAN performs well on classes\nwell-represente
 d during training, its accuracy and confidence\ndegrade substantially for 
 less common categories. Large language\nmodels (LLMs) offer a promising so
 lution for underrepresented\noncological entities, owing to their ability 
 to\nleverage context and pretraining. Rather than relying solely\non gener
 al-purpose models, domain adaptation or continual\npretraining of LLMs may
  further improve performance by\nhelping models learn the specialized voca
 bulary, abbreviations,\nand context typical of clinical text. In this stud
 y, we finetune\nLLMs for SEER pathology report classification, with and\nw
 ithout additional domain-adaptive pretraining, and compare\nthe results to
  the HiSAN baseline [2].\n\nBased on Llama 3 8B, PathLlama was developed b
 y finetuning\nfor cancer pathology report classification, with and without
 \ndomain adaptation. The domain adaptation task was next token\nprediction
  and the pretraining dataset was composed of a\nlarge corpus of approximat
 ely 10M cancer pathology reports\nand abstracts from SEER and about 500k c
 linical notes and\nradiology reports from MIMIC [3]. The PathLlama models\
 nwere finetuned to classify site (70 categories), subsite (330),\nlaterali
 ty (7), histology (677), and behavior (4). The finetuning\ndataset was 405
 2951 reports from six SEER registries:\nKentucky, Louisiana, New Jersey, N
 ew Mexico, Seattle/Puget\nSound, and Utah. The finetuning dataset was rand
 omly split\ninto 80%/10%/10% for training, test, and validation, ensuring\
 nall reports associated with a single case belong to the same\nsplit.\n\nF
 inetuning results are shown in Table I. We observe that\nthe micro F1 scor
 es, dominated by majority classes due the\nimbalance in the dataset, impro
 ve only slightly from the\nHiSAN to either of the PathLlama models. The mo
 st notable\nimprovements in micro F1 come from the domain-adapted\nPathLla
 ma for subsite and laterality. In contrast, more significant\nimprovements
  occur for macro F1, particularly for subsite,\nlaterality, and histology.
  For these three tasks, the domain-adapted\nPathLlama model also substanti
 ally outperforms the\nPathLlama base model. From these macro F1 results, w
 e find\nthat the contextual and pretraining advantages of Llama itself\nar
 e indeed sufficient to markedly improve classification performance\non und
 errepresented classes. However, domain adaptation\noffers additional benef
 it, further enhancing performance\nthat justifies the increased computatio
 nal cost associated with\nextended pretraining.\n\nRecording: Livestreamed
 , Recorded\n\nRegistration Category: Technical Program Reg Pass, Workshop 
 Reg Pass\n\nSession Chairs: Eric Stahlberg (MD Anderson Cancer Center, Uni
 versity of Texas); Sally Ellingson (University of Kentucky); Lynn Borkon (
 Frederick National Laboratory for Cancer Research); Patricia Kovatch (Icah
 n School of Medicine at Mount Sinai); Lauren Lewis (Frederick National Lab
 oratory for Cancer Research); and Sean Hanlon (National Institutes of Heal
 th (NIH), National Cancer Institute (NCI))\n\n
END:VEVENT
END:VCALENDAR
