Engineering PapersSearch

SEARCH · Engineering Papers

Results for “cancer registry”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence

Evaluating algorithmic bias on biomarker classification of breast cancer pathology reports

Objectives: This work evaluated algorithmic bias in biomarkers classification using electronic pathology reports from female breast cancer cases. Bias was assessed across 5 subgroups: cancer registry, race, Hispanic ethnicity, age at diagnosis, and socioeconomic status. Materials and Methods: We utilized 594 875 electronic pathology reports from 178 121 tumors diagnosed in Kentucky, Louisiana, New Jersey, New Mexico, Seattle, and Utah to train 2 deep-learning algorithms to classify breast cancer patients using their biomarkers test results. We used balanced error rate (BER), demographic parity (DP), equalized odds (EOD), and equal opportunity (EOP) to assess bias. Results: We found differences in predictive accuracy between registries, with the highest accuracy in the registry that contributed the most data (Seattle Registry, BER ratios for all registries >1.25). BER showed no significant algorithmic bias in extracting biomarkers (estrogen receptor, progesterone receptor, human epidermal growth factor receptor 2) for race, Hispanic ethnicity, age at diagnosis, or socioeconomic subgroups (BER ratio <1.25). DP, EOD, and EOP all showed insignificant results. Discussion: We observed significant differences in BER by registry, but no significant bias using the DP, EOD, and EOP metrics for socio-demographic or racial categories. This highlights the importance of employing a diverse set of metrics for a comprehensive evaluation of model fairness. Conclusion: A thorough evaluation of algorithmic biases that may affect equality in clinical care is a critical step before deploying algorithms in the real world. We found little evidence of algorithmic bias in our biomarker classification tool. Artificial intelligence tools to expedite information extraction from clinical records could accelerate clinical trial matching and improve care.

60 APPLIED LIFE SCIENCES

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)

Global Explainability of A Deep Abstaining Classifier for Cancer Pathology Reports

We present a global explainability method to characterize sources of errors in a real-world multitask deep abstaining classifier (DAC), in the context of cancer histology prediction. Our multitask classifier, currently deployed for automated annotation of cancer pathology reports from NCI-SEER registries, was trained and evaluated on 1.04 million hand-annotated samples and makes simultaneous predictions of cancer site, subsite, histology, laterality, and behavior for each report. The DAC framework enables the model to abstain on ambiguous reports and confusing classes to achieve the target accuracy on the retained (non-abstained) samples, but at the cost of decreased coverage. Requiring 97% accuracy on the histology task caused our model to retain only 22% of all samples, mostly the less ambiguous and common classes. Local explainability with the GradInp technique provided a computationally efficient way of obtaining contextual reasoning for hundreds of thousands of individual predictions. Our method, involving dimensionality reduction of approximately 13000 aggregated local explanations (ALE), offers a tractable path to true global explainability. It enabled identification of sources of errors in histology classification, globally, as hierarchical complexity among classes, label noise, insufficient information, and conflicting evidence. This suggests several strategies for iterative improvement of our DAC, including well-designed exclusion criteria, focused annotation, and reduced penalties for errors involving hierarchically related classes.

59 BASIC BIOLOGICAL SCIENCES

National Cancer Institute (NCI) Exposomic Linkage Protocol

The purpose of this work is to create point-level linkages of residential history data, which is provided by the Surveillance, Epidemiology, and End Results (SEER) program, to air pollution exposure data so that we can develop longitudinal measures of exposure and investigate their effects on cancer incidence, treatment response, and survival. The Louisiana, New Jersey, Kentucky, and Iowa registries were previously linked to LexisNexis residential history data through the National Cancer Institute (NCI). We will enhance the utility of the existing residential location data by geocoding addresses based on data from between 1995 and 2024 and spatially linking the locations to air pollution, indoor radon, and the US Environmental Protection Agency’s (EPA) Risk-Screening Environmental Indicators (RSEI) exposure data (Figure 1).

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN

Large-scale deep learning for metastasis detection in pathology reports

Objectives No existing algorithm can reliably identify metastasis from pathology reports across multiple cancer types and the entire US population. In this study, we develop a deep learning model that automatically detects patients with metastatic cancer by using pathology reports from many laboratories and of multiple cancer types. Materials and Methods We use 60 471 unstructured pathology reports from 4 Surveillance, Epidemiology, and End Results (SEER) registries. The reports were coded into 1 of 3 labels: metastasis negative, metastases positive, or metastasis undetermined. We utilize a task-specific deep neural network trained from scratch and compare its performance with a widely used large language model (LLM). Results Our deep learning architecture trained on task-specific data outperforms a general-purpose LLM, with a recall of 0.894 compared to 0.824. We quantified model uncertainty and used it to defer reports for human review. We found that retaining 72.9% of reports increased recall from 0.894 to 0.969. Discussion A smaller deep learning architecture trained on task-specific data outperforms a general LLM. Equally critical to model performance is the incorporation of uncertainty quantification, achieved here through an abstention mechanism. Conclusions This study’s finding demonstrate the feasibility of developing algorithms to automatically identify metastatic cancer cases from unstructured pathology reports.

machine learning

Archival records housed at USTUR support radium dial worker dosimetry

The American radium dial worker (RDW) cohort of over 3200 persons is being revisited as part of the Million Person Study (MPS) to include a modern approach to RDW dosimetry. An exceptional source of data and contextualization in this project is an extensive collection of electronic records (digitized from existing microfilm and microfiche) housed at the United States Transuranium and Uranium Registries (USTUR). Although the type, extent, and quality (e.g. legibility) of record(s) varies between individuals, the remarkable occupational, medical and demographic data include in vivo radiation measurements (e.g. radon breath, whole body counts), autopsy results, medical records (including copies of radiographs), interviews over the years, and correspondence. Of particular dosimetric interest are the details of radiation measurements. For example, there are some instances where hand-written and transcribed values are both available, along with notes providing context for why a particular measurement in a series of measurements was chosen to assign an intake, or if there were concerns about a particular measurement. Born prior to 1935, RDW have nearly all passed away. Thus, the updated dosimetry, especially for the skeletal tissues, will allow the correlation of lifetime cumulative dose with radiation risk. Here we review typical information available in this collection of historical records and highlight some interesting finds. Additionally, we discuss the relevance to current and ongoing work related to updating the dosimetry of the RDW in the MPS, including providing an example of the usefulness of information contained in these records. The RDW cohort provides a unique historical perspective on occupational exposure to radium, making it a valuable dataset for understanding long-term health effects and improving current radiation protection standards.

Million Person Study

Distribution of plutonium and radium in the human heart

Since 1968, the United States Transuranium and Uranium Registries (USTUR) has studied the biokinetics and tissue dosimetry of uranium and transuranium elements in nuclear workers. As part of the USTUR collaboration with the Million Person Study of Low-Dose Health Effects, radiation dose to different parts of the human heart is being estimated for workers with documented intakes of 239 Pu or 226 Ra. The study may be expanded for workers with intakes of 238 U and other radionuclides. The distribution of radionuclides, expressed in terms of concentration (Bq per kg of tissue) serves as an important parameter for estimating radiation dose. Based on available organs from workers who donated their bodies or tissues for research, nine undissected hearts were selected: seven from USTUR registrants with plutonium exposure (males) and two individuals with radium intakes (female and male). For the plutonium workers, estimated 239 Pu systemic deposition ranged from <74 Bq to 1765 Bq. Estimated 226 Ra ‘initial systemic intakes’ were 10.1 MBq and 14.8 kBq for the female patient and male worker, respectively. Organ dissection was based on a heart model published by Borrego et al (2019 J. Radiol. Prot. 39 950–65). This model includes nine cardiac substructures: aorta, left main coronary artery, left atrium, left anterior descending artery, left circumflex artery, left ventricle, right atrium, right coronary artery, and right ventricle. In addition, heart valves, fat attached to epicardium, fluids, and a coronary bypass graft were collected resulting in 111 samples that are currently undergoing radiochemical analyses and mass-spectrometric measurements. The 239 Pu and 226 Ra evaluations are not completed. The results of this study are intended to support radiation worker health studies by improving associated dosimetric and epidemiological models.

USTUR