Engineering Papers⌕ Search

Engineering topics

Hanson, Heidi A.

Publications and source records attributed to Hanson, Heidi A..

Landscape analysis of environmental data sources for linkage with SEER cancer patients database

Abstract One of the challenges associated with understanding environmental impacts on cancer risk and outcomes is estimating potential exposures of individuals diagnosed with cancer to adverse environmental conditions over the life course. Historically, this has been partly due to the lack of reliable measures of cancer patients’ potential environmental exposures before a cancer diagnosis. The emerging sources of cancer-related spatiotemporal environmental data and residential history information, coupled with novel technologies for data extraction and linkage, present an opportunity to integrate these data into the existing cancer surveillance data infrastructure, thereby facilitating more comprehensive assessment of cancer risk and outcomes. In this paper, we performed a landscape analysis of the available environmental data sources that could be linked to historical residential address information of cancer patients’ records collected by the National Cancer Institute’s Surveillance, Epidemiology, and End Results Program. The objective is to enable researchers to use these data to assess potential exposures at the time of cancer initiation through the time of diagnosis and even after diagnosis. The paper addresses the challenges associated with data collection and completeness at various spatial and temporal scales, as well as opportunities and directions for future research.

60 APPLIED LIFE SCIENCES↗

Urban relatives ameliorate survival disparities for genitourinary cancer in rural patients

Patients living in rural areas have worse cancer-specific outcomes. This study examines the effect of family-based social capital on genitourinary cancer survival. We hypothesized that rural patients with urban relatives have improved survival relative to rural patients without urban family. We examined rural and urban based Utah individuals diagnosed with genitourinary cancers between 1968 and 2018. Familial networks were determined using the Utah Population Database. Patients and relatives were classified as rural or urban based on 2010 rural–urban commuting area codes. Overall survival was analyzed using Cox proportional hazards models. We identified 24,746 patients with genitourinary cancer with a median follow-up of 8.72 years. Rural cancer patients without an urban relative had the worst outcomes with cancer-specific survival hazard ratios (HRs) at 5 and 10 years of 1.33 (95% CI 1.10–1.62) and 1.46 (95% CI 1.24–1.73), respectively relative to urban patients. Rural patients with urban first-degree relatives had improved survival with 5- and 10-year survival HRs of 1.21 (95% CI 1.06–1.40) and 1.16 (95% CI 1.03–1.31), respectively. Our findings suggest rural patients who have been diagnosed with a genitourinary cancer have improved survival when having relatives in urban centers relative to rural patients without urban relatives. Further research is needed to better understand the mechanisms through which having an urban family member contributes to improved cancer outcomes for rural patients. Better characterization of this affect may help inform policies to reduce urban–rural cancer disparities.

60 APPLIED LIFE SCIENCES↗

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗

Describing patterns of familial cancer risk in subfertile men using population pedigree data

STUDY QUESTION Can we simultaneously assess risk for multiple cancers to identify familial multicancer patterns in families of azoospermic and severely oligozoospermic men? SUMMARY ANSWER Here, distinct familial cancer patterns were observed in the azoospermia and severe oligozoospermia cohorts, suggesting heterogeneity in familial cancer risk by both type of subfertility and within subfertility type. WHAT IS KNOWN ALREADY Subfertile men and their relatives show increased risk for certain cancers including testicular, thyroid, and pediatric. STUDY DESIGN, SIZE, DURATION A retrospective cohort of subfertile men (N = 786) was identified and matched to fertile population controls (N = 5674). Family members out to third-degree relatives were identified for both subfertile men and fertile population controls (N = 337 754). The study period was 1966–2017. Individuals were censored at death or loss to follow-up, loss to follow-up occurred if they left Utah during the study period. PARTICIPANTS/MATERIALS, SETTING, METHODS Azoospermic (0 × 10 6 /mL) and severely oligozoospermic (<1.5 × 10 6 /mL) men were identified in the Subfertility Health and Assisted Reproduction and the Environment cohort (SHARE). Subfertile men were age- and sex-matched 5:1 to fertile population controls and family members out to third-degree relatives were identified using the Utah Population Database (UPDB). Cancer diagnoses were identified through the Utah Cancer Registry. Families containing ≥10 members with ≥1 year of follow-up 1966–2017 were included (azoospermic: N = 426 families, 21 361 individuals; oligozoospermic: N = 360 families, 18 818 individuals). Unsupervised clustering based on standardized incidence ratios for 34 cancer phenotypes in the families was used to identify familial multicancer patterns; azoospermia and severe oligospermia families were assessed separately. MAIN RESULTS AND THE ROLE OF CHANCE Compared to control families, significant increases in cancer risks were observed in the azoospermia cohort for five cancer types: bone and joint cancers hazard ratio (HR) = 2.56 (95% CI = 1.48–4.42), soft tissue cancers HR = 1.56 (95% CI = 1.01–2.39), uterine cancers HR = 1.27 (95% CI = 1.03–1.56), Hodgkin lymphomas HR = 1.60 (95% CI = 1.07–2.39), and thyroid cancer HR = 1.54 (95% CI = 1.21–1.97). Among severe oligozoospermia families, increased risk was seen for three cancer types: colon cancer HR = 1.16 (95% CI = 1.01–1.32), bone and joint cancers HR = 2.43 (95% CI = 1.30–4.54), and testis cancer HR = 2.34 (95% CI = 1.60–3.42) along with a significant decrease in esophageal cancer risk HR = 0.39 (95% CI = 0.16–0.97). Thirteen clusters of familial multicancer patterns were identified in families of azoospermic men, 66% of families in the azoospermia cohort showed population-level cancer risks, however, the remaining 12 clusters showed elevated risk for 2-7 cancer types. Several of the clusters with elevated cancer risks also showed increased odds of cancer diagnoses at young ages with six clusters showing increased odds of adolescent and young adult (AYA) diagnosis [odds ratio (OR) = 1.96–2.88] and two clusters showing increased odds of pediatric cancer diagnosis (OR = 3.64–12.63). Within the severe oligozoospermia cohort, 12 distinct familial multicancer clusters were identified. All 12 clusters showed elevated risk for 1–3 cancer types. An increase in odds of cancer diagnoses at young ages was also seen in five of the severe oligozoospermia familial multicancer clusters, three clusters showed increased odds of AYA diagnosis (OR = 2.19–2.78) with an additional two clusters showing increased odds of a pediatric diagnosis (OR = 3.84–9.32). LIMITATIONS, REASONS FOR CAUTION Although this study has many strengths, including population data for family structure, cancer diagnoses and subfertility, there are limitations. First, semen measures are not available for the sample of fertile men. Second, there is no information on medical comorbidities or lifestyle risk factors such as smoking status, BMI, or environmental exposures. Third, all of the subfertile men included in this study were seen at a fertility clinic for evaluation. These men were therefore a subset of the overall population experiencing fertility problems and likely represent those with the socioeconomic means for evaluation by a physician. WIDER IMPLICATIONS OF THE FINDINGS This analysis leveraged unique population-level data resources, SHARE and the UPDB, to describe novel multicancer clusters among the families of azoospermic and severely oligozoospermic men. Distinct overall multicancer risk and familial multicancer patterns were observed in the azoospermia and severe oligozoospermia cohorts, suggesting heterogeneity in cancer risk by type of subfertility and within subfertility type. Describing families with similar cancer risk patterns provides a new avenue to increase homogeneity for focused gene discovery and environmental risk factor studies. Such discoveries will lead to more accurate risk predictions and improved counseling for patients and their families. STUDY FUNDING/COMPETING INTEREST(S) This work was funded by GEMS: Genomic approach to connecting Elevated germline Mutation rates with male infertility and Somatic health (Eunice Kennedy Shriver National Institute of Child Health and Human Development (NICHD): R01 HD106112). The authors have no conflicts of interest relevant to this work.

60 APPLIED LIFE SCIENCES↗

Evaluating county-level lung cancer incidence from environmental radiation exposure, PM 2.5 , and other exposures with regression and machine learning models

Characterizing the interplay between exposures shaping the human exposome is vital for uncovering the etiology of complex diseases. For example, cancer risk is modified by a range of multifactorial external environmental exposures. Environmental, socioeconomic, and lifestyle factors all shape lung cancer risk. However, epidemiological studies of radon aimed at identifying populations at high risk for lung cancer often fail to consider multiple exposures simultaneously. For example, moderating factors, such as PM 2.5 , may affect the transport of radon progeny to lung tissue. This ecological analysis leveraged a population-level dataset from the National Cancer Institute’s Surveillance, Epidemiology, and End-Results data (2013–17) to simultaneously investigate the effect of multiple sources of low-dose radiation (gross γ activity and indoor radon) and PM 2.5 on lung cancer incidence rates in the USA. County-level factors (environmental, sociodemographic, lifestyle) were controlled for, and Poisson regression and random forest models were used to assess the association between radon exposure and lung and bronchus cancer incidence rates. Tree-based machine learning (ML) method perform better than traditional regression: Poisson regression: 6.29/7.13 (mean absolute percentage error, MAPE), 12.70/12.77 (root mean square error, RMSE); Poisson random forest regression: 1.22/1.16 (MAPE), 8.01/8.15 (RMSE). The effect of PM 2.5 increased with the concentration of environmental radon, thereby confirming findings from previous studies that investigated the possible synergistic effect of radon and PM 2.5 on health outcomes. In summary, the results demonstrated (1) a need to consider multiple environmental exposures when assessing radon exposure’s association with lung cancer risk, thereby highlighting (1) the importance of an exposomics framework and (2) that employing ML models may capture the complex interplay between environmental exposures and health, as in the case of indoor radon exposure and lung cancer incidence.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

FrESCO: Framework for Exploring Scalable Computational Oncology

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information from electronic health records. This data is then used to create summary statistics about cancer incidence and mortality to facilitate population health monitoring. Extracting phenotypic information from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating the information extraction process from cancer pathology reports has the potential to improve data quality by extracting information in a consistent manner across registries. It can also improve patient outcomes by reducing the time from diagnosis, enabling rapid case ascertainment for clinical trials. Here we present FrESCO, a modular deep-learning natural language processing (NLP) library initially designed for extracting pathology information from clinical text documents. This repository is not solely limited to clinical medical text, but may also be used by researchers just getting started with NLP methods and those looking for a robust solution for their classification problems.

60 APPLIED LIFE SCIENCES↗

Dataset Repository for Investigating Suicide Risk Using Social and Environmental Determinants of Health

Suicide is frequently modeled as a function of genetics and environment, where the latter refers to factors other than direct biological consequences, such as air quality, financial level, social connectivity, transportation and food access, and homelessness status. According to the World Health Organization, clean air, a stable climate, adequate water, sanitation and hygiene, safe chemical use, radiation protection, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved natural environment are all prerequisites for good health. Understanding the relationships between these determinants and mental health outcomes requires standardized data that can be included in healthcare programs and health outcome models. There is a wealth of publicly available data on social and environmental factors provided by various US organizations that can benefit the design of health care systems and public health interventions, as well as improve our comprehension of factors that impact health. Such information would not only help improve the understanding of individual and community risk but also identify new risk factors that have not previously been therapeutically targeted, especially in terms of their impact on mental health. However, curating and standardizing such datasets is challenging because they are often recorded at numerous geographical and temporal resolutions and with varying spatial and temporal granularities. To address this challenge, we launched an endeavor in conjunction with the Veterans Health Administration to collect publicly available socioeconomic and environmental determinants of health statistics in the US. In this manuscript, we describe a social and environmental determinants of health (SEDH) datasets repository, data curation documentation, and a pipeline framework for data generation; This effort started in 2020, when we began constructing a scalable pipeline to automate the download, extraction, preparation, analysis, and production of datasets. These datasets have been made available to the VHA and may be shared upon agreement with collaborating organizations.

60 APPLIED LIFE SCIENCES↗

Familial Associations of Prevalence and Cause-Specific Mortality for Thoracic Aortic Disease and Bicuspid Aortic Valve in a Large-Population Database

Thoracic aortic disease and bicuspid aortic valve (BAV) likely have a heritable component, but large population-based studies are lacking. This study characterizes familial associations of thoracic aortic disease and BAV, as well as cardiovascular and aortic-specific mortality, among relatives of these individuals in a large-population database. In this observational case-control study of the Utah Population Database, we identified probands with a diagnosis of BAV, thoracic aortic aneurysm, or thoracic aortic dissection. Age- and sex-matched controls (10:1 ratio) were identified for each proband. First-degree relatives, second-degree relatives, and first cousins of probands and controls were identified through linked genealogical information. Cox proportional hazard models were used to quantify the familial associations for each diagnosis. We used a competing-risk model to determine the risk of cardiovascular-specific and aortic-specific mortality for relatives of probands. The study population included 3,812,588 unique individuals. Familial hazard risk of a concordant diagnosis was elevated in the following populations compared with controls: first-degree relatives of patients with BAV (hazard ratio [HR], 6.88 [95% CI, 5.62–8.43]); first-degree relatives of patients with thoracic aortic aneurysm (HR, 5.09 [95% CI, 3.80–6.82]); and first-degree relatives of patients with thoracic aortic dissection (HR, 4.15 [95% CI, 3.25–5.31]). In addition, the risk of aortic dissection was higher in first-degree relatives of patients with BAV (HR, 3.63 [95% CI, 2.68–4.91]) and in first-degree relatives of patients with thoracic aneurysm (HR, 3.89 [95% CI, 2.93–5.18]) compared with controls. Dissection risk was highest in first-degree relatives of patients who carried a diagnosis of both BAV and aneurysm (HR, 6.13 [95% CI, 2.82–13.33]). First-degree relatives of patients with BAV, thoracic aneurysm, or aortic dissection had a higher risk of aortic-specific mortality (HR, 2.83 [95% CI, 2.44–3.29]) compared with controls. Our results indicate that BAV and thoracic aortic disease carry a significant familial association for concordant disease and aortic dissection. The pattern of familiality is consistent with a genetic cause of disease. Furthermore, we observed higher risk of aortic-specific mortality in relatives of individuals with these diagnoses. In conclusion, this study provides supportive evidence for screening in relatives of patients with BAV, thoracic aneurysm, or dissection.

60 APPLIED LIFE SCIENCES↗

Environmental exposure to industrial air pollution is associated with decreased male fertility

Objective: To understand how chronic exposure to industrial air pollution is associated with male fertility through semen parameters. Design: Retrospective cohort study. Subjects: Men in the Subfertility, Health and Assisted Reproduction cohort who underwent a semen analysis 2005-2017 with ≥1 measured semen parameter (N=21,563). Intervention(s): Residential histories for each man were constructed using locations from administrative records linked through the Utah Population Database. Industrial facilities with air emissions of nine endocrine disrupting compound chemical classes were identified from the Environmental Protection Agency Risk-Screening Environmental Indicators microdata. Chemical levels were linked with residential histories for the 5 years prior to each semen analysis. Main Outcome Measures: Semen analyses were classified as azoospermic or oligozoospermic (< 15 M/mL) using World Health Organization cutoffs for concentration. Bulk semen parameters such as concentration, total count, ejaculate volume, total motility, total motile count, and total progressive motile count were also measured. Multivariable regression models with robust standard errors were used to associate exposure quartiles for each of the nine chemical classes with each semen parameter, adjusting for age, race, and ethnicity, as well as neighborhood socioeconomic disadvantage. Results: After adjustment for demographic covariates, several chemical classes were associated with azoospermia and decreased total motility and volume. For exposure in the 4th relative to 1st quartile, significant associations were observed for acrylonitrile (β total motility = -0.87 pp), aromatic hydrocarbons (odds ratio [OR]azoospermia = 1.53; β volume = -0.14 mL), dioxins (OR azoospermia = 1.31; β volume = -0.09 mL; β total motility = -2.65 pp), heavy metals (β total motility = -2.78pp), organic solvents (OR azoospermia = 1.75; β volume = -0.10 mL), organochlorines (OR azoospermia = 2.09; β volume = -0.12 mL), phthalates (OR azoospermia = 1.44; β volume = -0.09 mL; β total motility = -1.21 pp), and silver particles (OR azoospermia = 1.64; β volume = -0.11 mL). All semen parameters significantly decreased with increasing socioeconomic disadvantage. Men who lived in the most disadvantaged areas had concentration, volume, and total motility of 6.70 M/mL, 0.13 mL, and 1.79 pp lower, respectively. Count, motile count, and total progressive motile count all decreased by 30–34 M.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Risk Factors and Trends for HPV-Associated Subsequent Malignant Neoplasms among Adolescent and Young Adult Cancer Survivors

Subsequent malignant neoplasms (SMN; new cancers that arise after an original diagnosis) contribute to premature mortality among adolescent and young adult (AYA) cancer survivors. Because of the high population prevalence of human papillomavirus (HPV) infection, we identify demographic and clinical risk factors for HPV-associated SMNs (HPV-SMN) among AYA cancer survivors in the SEER-9 registries diagnosed from 1976 to 2015. Outcomes included any HPV-SMN, oropharyngeal-SMN, and cervical-SMN. Follow-up started 2 months after their original diagnosis. Standardized incidence ratios (SIR) compared risk between AYA survivors and general population. Age-period-cohort (APC) models examined trends over time. Fine and Gray's models identified therapy effects controlling for cancer and demographic confounders. Of 374,408 survivors, 1,369 had an HPV-SMN, occurring on average 5 years after first cancer. Compared with the general population, AYA survivors had 70% increased risk for any HPV-SMN [95% confidence interval (CI), 1.61–1.79] and 117% for oropharyngeal-SMN (95% CI, 2.00–2.35); cervical-SMN risk was generally lower in survivors (SIR, 0.85; 95% CI, 0.76–0.95), but Hispanic AYA survivors had a 8.4 significant increase in cervical-SMN (SIR, 1.46; 95% CI, 1.01–2.06). AYAs first diagnosed with Kaposi sarcoma, leukemia, Hodgkin, and non-Hodgkin lymphoma had increased HPV-SMN risks compared with the general population. Oropharyngeal-SMN incidence declined over time in APC models. Chemotherapy and radiation were associated with any HPV-SMN among survivors with first HPV-related cancers, but not associated among survivors whose first cancers were not HPV-related. HPV-SMN in AYA survivors are driven by oropharyngeal cancers despite temporal declines in oropharyngeal-SMN. Hispanic survivors are at risk for cervical-SMN relative to the general population. Encouraging HPV vaccination and cervical and oral cancer screenings may reduce HPV-SMN burden among AYA survivors.

60 APPLIED LIFE SCIENCES↗

Generating Older Adult Multimorbidity Trajectories Using Various Comorbidity Indices and Calculation Methods

Older adult multimorbidity trajectories are helpful for understanding the current and future health patterns of aging populations. The construction of multimorbidity trajectories from comorbidity index scores will help inform public health and clinical interventions targeting those individuals that are on unhealthy trajectories. Investigators have used many different techniques when creating multimorbidity trajectories in prior literature, and no standard way has emerged. This study compares and contrasts multimorbidity trajectories constructed from various methods. We describe the difference between aging trajectories constructed with the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index (ECI). We also explore the differences between acute (single year) and chronic (cumulative) derivations of CCI and ECI scores. Social determinants of health can affect disease burden over time; thus, our models include income, race/ethnicity, and sex differences. We use group-based trajectory modeling (GBTM) to estimate multimorbidity trajectories for 86,909 individuals aged 66–75 in 1992 using Medicare claims data collected over the following 21 years. We identify low-chronic disease and high-chronic disease trajectories in all 8 generated trajectory models. Additionally, all 8 models satisfied prior established statistical diagnostic criteria for well-performing GBTM models. Clinicians may use these trajectories to identify patients on an unhealthy path and prompt a possible intervention that may shift the patient to a healthier trajectory.

60 APPLIED LIFE SCIENCES↗