Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “human reliability analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Supply Chain Risk Management: Data Structuring

Supply chain risk management (SCRM) is an area of research that addresses both logistics concepts to maximize efficiency, reliability, and revenue as well as risk features, such as potential weak points, break points, and vulnerabilities within the supply chain. SCRM is used to find risks introduced at each node in a supply chain and how these risks can impact a company’s products, individuals, customers, and reputation. SCRM is a relatively new field, so standardized processes including data structuring are not fully documented. This paper explains the importance of a standard data structuring methodology and how it can enhance current SCRM efforts. Data ingest, structuring, and analysis are predominantly managed by humans. Automating some of the less complex steps can positively impact SCRM by allowing human analysts to focus on more strategic analyses. Types of data to be collected and structured are collected via publicly available information related to hardware, software, and corporate entities. After the data has been collected, the information is formatted in a specific manner, conforming to a schema, to allow for more effective and efficient ingest for further analysis. This paper outlines data structures used by Pacific Northwest National Laboratory for SCRM research and analysis purposes. These structures have been used for hundreds of analyses and have been successful in developing a common baseline. Data structuring is one of the first steps in data standardization, which will further mature and enhance the SCRM research area.

supply chain risk management, data structuring, re↗

Generalized analytical and numerical modeling of optical second harmonic generation in anisotropic crystals and complex heterostructures using #SHAARP package

Optical second harmonic generation (SHG) is a nonlinear optical effect widely used for nonlinear optical microscopy and laser frequency conversion. The closed-form analytical solution of the nonlinear optical responses is essential for evaluating the optical responses of new materials whose optical properties are unknown a priori. Many approximations have therefore been employed in the existing analytical approaches, such as slowly varying approximation, weak reflection of the nonlinear polarization, transparent medium, high crystallographic symmetry, Kleinman symmetry, easy crystal orientation along a high-symmetry direction, phase matching conditions and negligible interference among nonlinear waves, which may lead to large errors in the reported material properties. To avoid these approximations, here we have developed an open-source package named Second Harmonic Analysis of Anisotropic Rotational Polarimetry (#SHAARP) for single interface (si) and in multilayers (ml) for homogeneous crystals. The reliability and accuracy are established by experimentally benchmarking with both the SHG polarimetry and Maker fringes predicted from the package using standard materials. SHAARP.si and SHAARP.ml are available through GitHub https://github.com/Rui-Zu/SHAARP and https://github.com/bzw133/SHAARP.ml, respectively.

complex systems↗

Agentic Diagrammatica: Towards Autonomous Symbolic Computation in High Energy Physics

We present Diagrammatica, a symbolic computation extension to the HEPTAPOD agentic framework, which enables LLM agents to plan and execute multi-step theoretical calculations. Symbolic computation poses a distinctive reliability challenge for LLM agents, as correctness is governed by implicit mathematical conventions that are not encoded in a form that can be easily checked in the computational backend. We identify two complementary remedies, tool-constrained computation and targeted knowledge grounding, and pursue the first as the primary architecture. Concretely, we concentrate the agent's action distribution onto tool calls with convention-fixing semantics, in which the agent specifies a compact, human-auditable diagram specification and a trusted backend performs the symbolic or numerical manipulations exactly. The toolkit provides two complementary calculation paths consuming a shared diagram specification: Naive Dimensional Analysis (NDA) for order-of-magnitude rate estimates and Exact Diagrammatic Analysis (EDA) for tree-level symbolic calculations via automatic FeynCalc code generation, both supplemented by automatic Feynman diagram enumeration and a navigable theory knowledge base. The architecture is validated on two benchmarks: (1) an exhaustive catalog of all tree-level, single-vertex $1\to 2$ partial decay widths across scalar, fermion, and vector parents, with complete massless and threshold limits and Standard Model validation; and (2) an NDA sensitivity study of the muon decay multiplicity $μ^+ \to ν_μ\barν_e + n(e^+e^-) + e^-$, determining the maximum observable $n$ at current and planned muon experiments.

Menzo, Tony [Alabama U.; Fermilab] (ORCID:00000002↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Separation of Lipoproteins for Quantitative Analysis of 14 C-Labeled Lipid-Soluble Compounds by Accelerator Mass Spectrometry

To date, 14 C tracer studies using accelerator mass spectrometry (AMS) have not yet resolved lipid-soluble analytes into individual lipoprotein density subclasses. The objective of this work was to develop a reliable method for lipoprotein separation and quantitative recovery for biokinetic modeling purposes. The novel method developed provides the means for use of small volumes (10–200 µL) of frozen plasma as a starting material for continuous isopycnic lipoprotein separation within a carbon- and pH-stable analyte matrix, which, following post-separation fraction clean up, created samples suitable for highly accurate 14 C/ 12 C isotope ratio determinations by AMS. Manual aspiration achieved 99.2 ± 0.41% recovery of [5- 14 CH 3 ]-(2R, 4'R, 8'R)-α-tocopherol contained within 25 µL plasma recovered in triacylglycerol rich lipoproteins (TRL = Chylomicrons + VLDL), LDL, HDL, and infranatant (INF) from each of 10 different sampling times for one male and one female subject, n = 20 total samples. Small sample volumes of previously frozen plasma and high analyte recoveries make this an attractive method for AMS studies using newer, smaller footprint AMS equipment to develop genuine tracer analyses of lipophilic nutrients or compounds in all human age ranges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Flame stability analysis of flame spray pyrolysis by artificial intelligence

Flame spray pyrolysis (FSP) is a process used to synthesize nanoparticles through the combustion of an atomized precursor solution; this process has applications in catalysts, battery materials, and pigments. Current limitations revolve around understanding how to consistently achieve a stable flame and the reliable production of nanoparticles. Machine learning and artificial intelligence algorithms that detect unstable flame conditions in real time may be a means of streamlining the synthesis process and improving FSP efficiency. In this study, the FSP flame stability is first quantified by analyzing the brightness of the flame's anchor point. This analysis is then used to label data for both unsupervised and supervised machine learning approaches. The unsupervised learning approach allows for autonomous labeling and classification of new data by representing data in a reduced dimensional space and identifying combinations of features that most effectively cluster it. The supervised learning approach, on the other hand, requires human labeling of training and test data but is able to classify multiple objects of interest (such as the burner and pilot flames) within the video feed. The accuracy of each of these techniques is compared against the evaluations of human experts. Both the unsupervised and supervised approaches can track and classify FSP flame conditions in real time to alert users of unstable flame conditions. This research has the potential to autonomously track and manage flame spray pyrolysis as well as other flame technologies by monitoring and classifying the flame stability.

42 ENGINEERING↗

PSA 2025 DPRA for Cyber Optimization

Cyberattacks can have many different attack paths, durations, and goals. There are also many different mitigation options involving hardware, software, and/or humans. Evaluating defense options should include quantitative evaluation of overall effectiveness to make cost and risk-informed decisions. Typical cyberattack modeling methods only provide a qualitative evaluation and have difficulty with time dependent scenarios. The main areas of cybersecurity are confidentiality, integrity, and availability. For companies with cyber-physical systems such as advanced nuclear reactors, cyber-related integrity is a requirement set by the U.S. Nuclear Regulatory Commission. But companies are also concerned about availability or reliability as a business case. As cyber threats are evolving to a business-for-hire structure, more attacks focus on disrupting business success and reliability, causing financial and economic stability risk. Companies want reliability analysis while optimizing cost, which requires more than safety modeling methods. Dynamic-state-based and Markov-based modeling provides a method for better cyber scenario modeling with timing and conditional features not found in other numerical evaluation methods. EMRALD (Event Modeling Risk Assessment using Lined Diagrams) is a dynamic risk analysis modeling and simulation tool and has features that reduce modeling issues such as state-base explosion found in Markov-based tools. It has been used to model different time-dependent events including plant behavior and operator procedures. As a general modeling tool, EMRALD can also be used to model cyberattack scenarios with varying mitigation options and quantify effectiveness, producing numerical data for risk-informed decisions. This paper uses EMRALD to demonstrate that dynamic risk analysis can be used for cyber threat modeling to provide insights for design decision-making and optimize defense strategies.

97 - MATHEMATICS AND COMPUTING↗

Deep learning uncertainty quantification for clinical text classification

Machine learning algorithms are expected to work side-by-side with humans in decision-making pipelines. Thus, the ability of classifiers to make reliable decisions is of paramount importance. Deep neural networks (DNNs) represent the state-of-the-art models to address real-world classification. Although the strength of activation in DNNs is often correlated with the network’s confidence, in-depth analyses are needed to establish whether they are well calibrated. In this paper, we demonstrate the use of DNN-based classification tools to benefit cancer registries by automating information extraction of disease at diagnosis and at surgery from electronic text pathology reports from the US National Cancer Institute (NCI) Surveillance, Epidemiology, and End Results (SEER) population-based cancer registries. In particular, we introduce multiple methods for selective classification to achieve a target level of accuracy on multiple classification tasks while minimizing the rejection amount—that is, the number of electronic pathology reports for which the model’s predictions are unreliable. We evaluate the proposed methods by comparing our approach with the current in-house deep learning-based abstaining classifier. Overall, all the proposed selective classification methods effectively allow for achieving the targeted level of accuracy or higher in a trade-off analysis aimed to minimize the rejection rate. On in-distribution validation and holdout test data, with all the proposed methods, we achieve on all tasks the required target level of accuracy with a lower rejection rate than the deep abstaining classifier (DAC). Interpreting the results for the out-of-distribution test data is more complex; nevertheless, in this case as well, the rejection rate from the best among the proposed methods achieving 97% accuracy or higher is lower than the rejection rate based on the DAC. We show that although both approaches can flag those samples that should be manually reviewed and labeled by human annotators, the newly proposed methods retain a larger fraction and do so without retraining—thus offering a reduced computational cost compared with the in-house deep learning-based abstaining classifier.

59 BASIC BIOLOGICAL SCIENCES↗

Digital Infrastructure Migration Framework Report

This document presents a full-scope Digital Infrastructure implementation and associated lifecycle support recommendations that enable a plant life of 80+ years. Specific technologies and software applications are researched, developed, selected, implemented, and then integrated by utilities to enhance safety, reliability, and economic performance such that the result provides much more than the sum of its parts. Specific selection of these technologies is driven by business case analyses which are utility, station, and unit specific.

42 ENGINEERING↗

AI-Enabled Robots for Automated Nondestructive Evaluation and Repair of Power Plant Boilers. Final Report

Boiler failure could cause loss of life and safety issues, cost hundreds of thousands of dollars in equipment repairs, property damage and production losses, and drive up the cost of electric power. Boiler maintenance is challenging and risky for inspectors working on scaffolding in confined hazardous spaces inside of a boiler and sometimes the space is hard to access. The operation is also time-consuming due to the large area of vertical structures for inspection and the tremendous effort needed for scaffolding. Recently, the use of robotics (e.g., drones and crawlers) in power plants for maintenance is growing rapidly. However, the existing robotics solutions show two notable technological gaps: no live repair capability, and no Artificial Intelligence (AI) for smart autonomy. The objective of this project is to develop an integrated autonomous robotic platform that is equipped with compact non-destructive evaluation (NDE) sensors to perform live inspection, operates onboard repair devices to perform live repair, and uses AI for intelligent data fusion and predictive analysis for automated and smart spatiotemporal inspection, analysis and repair of the furnace walls in coal-fired boilers. The approach to achieve the objective includes developing NDE sensors with signal processing techniques, designing and evaluating repair devices for robots based on fusion and solid-state technologies, and an autonomous robotic platform that can attach to and navigate on boiler furnace walls using magnetic drive tracks. The robot is also powered by AI to automate data gathering (e.g., 3D mapping and damage localization) and predictive analysis. This project has advanced the state-of-the-art by providing technological breakthroughs including compact NDE and repair tools for robots, AI capabilities for smart autonomy, and a robotic platform for automated boiler maintenance. This project has great potential to result in significant benefits including limiting or eliminating the need to send operators to assess difficult-to-access or hazardous areas, enabling automated live inspection and repair, avoiding time consuming scaffolding (especially for partial maintenance during unplanned outage), collecting comprehensive and well-organized data smartly, and avoiding or limiting the need for onsite or remote piloting technicians. The impacts can be tremendous in terms of the time and cost savings, reducing the risk for human operators, and increasing boiler reliability, usability, and efficiency. In addition, by developing the new technologies on the autonomous inspection and repair robot, by involving multiple undergraduate and graduate students working together with the faculty members on this project, and by generating knowledge and building up collaborations with industrial partners, this effort will significantly update the education capabilities, support long-term fundamental research, and maintain the leadership of Colorado School of Mines and Michigan State University in energy fields.

20 FOSSIL-FUELED POWER PLANTS↗

Alpha Decay Chains as Thermal Power Sources: Analysis and Applications for RTGs

Radioactive sources can provide power in remote and environmentally harsh locations such as the arctic or space. The generators powered by such sources are rugged and can withstand extreme temperatures, lack of sunlight, and require no human intervention for multiple years. Radioisotopes are used in thermoelectric generators to provide power at remote sites and deep in space. Isotopes like Pu-238, Cm-244, and Am-241 are used in these generators by NASA for power in space probes and spacecrafts. These power sources deliver a steady supply of energy over extended periods of time. Alpha particles created during decay do not travel far in a material. Their kinetic energy is transferred to heat that we can then convert into energy. Unlike beta and gamma decay, the slower-moving alpha particles stop in the material, making their energy available for use. Energy from these natural decay processes provides a reliable source of power. Spontaneous fission is rare and unreliable, and unlike induced fission processes, alpha decay occurs naturally and does not require external management or ignition. The ideal properties of an isotope for use as a power source depend upon the intended use. For use in an Arctic research base over a period of several years, but less than a decade, an isotope that provides high power output over a shorter lifespan may be the most suitable option. Whereas, for deep space missions where a consistent power source for decades or perhaps more than 100 years is needed that would require a very different isotope. One with a much longer half-life that would provide consistent power throughout that time and survive in that state in for these extended periods of time. These examples represent two extreme sides in terms of time frames. By analyzing the power produced by different radioactive decay processes over time, we can evaluate the suitability of various isotope decay chains for specific uses. Some unstable isotopes undergo a series of radioactive decays, transforming into different isotopes at each step and resulting in a stable isotope. The lists of isotopes in these decay processes are known as decay chains. Some of these chains, illustrated in the figures below, are currently being investigated for use in radioisotope thermoelectric generators (RTGs) designed for a range of operational durations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Evaluating various composite sampling modes for detecting pathogenic SARS-CoV-2 virus in raw sewage

Inadequate sampling approaches to wastewater analyses can introduce biases, leading to inaccurate results such as false negatives and significant over- or underestimation of average daily viral concentrations, due to the sporadic nature of viral input. To address this challenge, we conducted a field trial within the University of Tennessee residence halls, employing different composite sampling modes that encompassed different time intervals (1 h, 2 h, 4 h, 6 h, and 24 h) across various time windows (morning, afternoon, evening, and late-night). Our primary objective was to identify the optimal approach for generating representative composite samples of SARS-CoV-2 from raw wastewater. Utilizing reverse transcription-quantitative polymerase chain reaction, we quantified the levels of SARS-CoV-2 RNA and pepper mild mottle virus (PMMoV) RNA in raw sewage. Our findings consistently demonstrated that PMMoV RNA, an indicator virus of human fecal contamination in water environment, exhibited higher abundance and lower variability compared to pathogenic SARS-CoV-2 RNA. Significantly, both SARS-CoV-2 and PMMoV RNA exhibited greater variability in 1 h individual composite samples throughout the entire sampling period, contrasting with the stability observed in other time-based composite samples. Through a comprehensive analysis of various composite sampling modes using the Quade Nonparametric ANCOVA test with date, PMMoV concentration and site as covariates, we concluded that employing a composite sampler during a focused 6 h morning window for pathogenic SARS-CoV-2 RNA is a pragmatic and cost-effective strategy for achieving representative composite samples within a single day in wastewater-based epidemiology applications. This method has the potential to significantly enhance the accuracy and reliability of data collected at the community level, thereby contributing to more informed public health decision-making during a pandemic.

sampling timing↗

Three-dimensional phenotyping of peach tree-crown architecture utilizing terrestrial laser scanning

Tree training systems for temperate fruit have been developed throughout history by pomologists to improve light interception, fruit yield, and fruit quality. These training systems direct crown and branch growth to specific configurations. Quantifying crown architecture could aid the selection of trees that require less pruning or that naturally excel in specific growing/training system conditions. Regarding peaches [Prunus persica (L.) Batsch], access tools such as branching indices have been developed to characterize tree-crown architecture. However, the required branching data (BD) to develop these indices are difficult to collect. Traditionally, BD have been collected manually, but this process is tedious, time-consuming, and prone to human error. These barriers can be circumnavigated by utilizing terrestrial laser scanning (TLS) to obtain a digital twin of the real tree. TLS generates three-dimensional (3D) point clouds of the tree crown, wherein every point contains 3D coordinates (x, y, z). To facilitate the use of these tools for peach, we selected 16 young peach trees scanned in 2021 and 2022. These 16 trees were then modeled and quantified using the open-source software TreeQSM. As a result, “in silico” branching and biometric data for the young peach trees were calculated to demonstrate the capabilities of TLS phenotyping of peach tree-crown architecture. The comparison and analysis of field measurements (in situ) and in silico BD, biometric data, and quantitative structural model branch uncertainty data were utilized to determine the reconstructive model’s reliability as a source substitute for field measurements. Mean average deviation when comparing young tree (YT) height was approx. 5.93%, with crown volume was approx. 13.26% across both 2021 and 2022. All point clouds of the YTs in 2022 showed residuals lower than 12 mm to cylinders fitted to all branches, and mean surface coverage greater than 40% for both the trunk and primary branching orders.

09 BIOMASS FUELS↗

Digitalization Guiding Principles and Method for Nuclear Industry Work Processes

The commercial U.S. light-water reactor fleet has been operating at historical efficiency, reliability, and safety over the last decade. Nuclear power has the highest capacity factor of any other power generation technology while also serving as the largest baseload source for carbon-free energy. Despite this remarkable achievement, continued operations for many plants are threatened due to fierce electricity market competition and rising operations and maintenance costs of which continued maintenance of obsolete analog equipment is a contributor. The digital age and associated technologies are where the future lies in process control, and nuclear has yet to take full advantage of the capabilities offered therein. The Light Water Reactor Sustainability Program (LWRS) at Idaho National Laboratory (INL), sponsored by the Department of Energy, has a mission to help the light-water reactor fleet manage its foundational capabilities to continue providing safe and reliable carbon-free power. LWRS helps support that mission by providing scientific, technology-based solutions for advanced concepts of operations with a more viable business model that will allow the fleet to continue to operate at peak levels through extended plant operation. The LWRS Digitalization Project at INL seeks to leverage digital technologies to synthesize and transform work processes. We provide a state-of-the-art analysis of digitalized work processes in nuclear power and investigate ways in which researchers at INL and the nuclear industry can work together to identify what data to access, how to access it, what to do with the data, and most importantly, how to use the insights for decision-making across all levels within the business. Borne from these considerations, we present four guiding principles for digitalization: develop a coherent digitalization plan, apply human factors engineering, establish data governance, and anticipate unintended consequences. Together, these principles form a method that plants can use to effectively to digitalize nuclear industry work processes. Our guiding principles are informed by multiple knowledge sources. First, we document activities from the Work Digitalization Initiative, which was conceived as a means for nuclear organizations to help define and standardize the industry’s approach to digitalizing work. Second, we detail primary research conducted with industry professionals regarding drivers and barriers to digitalization adoption. We present survey results that demonstrate what the industry hopes to get out of digitalization and the ways that INL can continue to support the industry’s digital transformation. Third, we present a digitalization use case with industry partners NextAxiom Technology and Xcel Energy. The project objective was to transform the current condition report work process from paper to digital, incorporating digitalized principles. We report the development of the application and lessons learned. The accomplishments achieved by this research and development serve to identify critical needs for plant guidance in support of digitalization implementation and contribute to the knowledge and strategies available for utilities considering or undertaking digitalization.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

In search of autophagy biomarkers in breast cancer: Receptor status and drug agnostic transcriptional changes during autophagy flux in cell lines

Autophagy drives drug resistance and drug-induced cancer cell cytotoxicity. Targeting the autophagy process could greatly improve chemotherapy outcomes. The discovery of specific inhibitors or activators has been hindered by challenges with reliably measuring autophagy levels in a clinical setting. We investigated drug-induced autophagy in breast cancer cell lines with differing ER/PR/Her2 receptor status by exposing them to known but divergent autophagy inducers each with a unique molecular target, tamoxifen, trastuzumab, bortezomib or rapamycin. Differential gene expression analysis from total RNA extracted during the earliest sign of autophagy flux showed both cell- and drug-specific changes. We analyzed the list of differentially expressed genes to find a common, cell- and drug-agnostic autophagy signature. Twelve mRNAs were significantly modulated by all the drugs and 11 were orthogonally verified with Q-RT-PCR (Klhl24, Hbp1, Crebrf, Ypel2, Fbxo32, Gdf15, Cdc25a, Ddit4, Psat1, Cd22, Ypel3). The drug agnostic mRNA signature was similarly induced by a mitochondrially targeted agent, MitoQ. In-silico analysis on the KM-plotter cancer database showed that the levels of these mRNAs are detectable in human samples and associated with breast cancer prognosis outcomes of Relapse-Free Survival in all patients (RSF), Overall Survival in all patients (OS), and Relapse-Free Survival in ER + Patients (RSF ER + ). High levels of Klhl24, Hbp1, Crebrf, Ypel2, CD22 and Ypel3 were correlated with better outcomes, whereas lower levels of Gdf15, Cdc25a, Ddit4 and Psat1 were associated with better prognosis in breast cancer patients. This gene signature uncovers candidate autophagy biomarkers that could be tested during preclinical and clinical studies to monitor the autophagy process.

60 APPLIED LIFE SCIENCES↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

BRAVE_EBC-TMT.1.0

Exhaled breath condensate (EBC) represents a low-cost and non-invasive means of examining respiratory health. EBC has been used to discover and validate exhaled volatile and non-volatile biomarkers of disease related to the respiratory system distress such as asthma, COPD, lung cancer, and secondary infections. One newly emerging utilization of EBC, is proteomics analysis, which can provide an unbiased snapshot into ongoing biological processes in the airway. Fully characterizing the biological landscape of EBC collections is challenging though, due to sample variability, and low detection sensitivity. EBC is primarily composed of condensed water, causing technical challenges with detecting key macromolecules from the dilute sample matrix; therefore, high sensitivity techniques are required to unlock the full capability of EBC as a method for non-invasive biomarker detection. To overcome some of these technical challenges for proteomic analyses, we applied our recently developed microscale proteomic techniques and developed a novel TMT based approach which enabled reliable, relative quantification with significantly improved detection of low abundance peptides/proteins across multiple healthy volunteer EBC samples. Our EBC collection design includes longitudinal EBC collections from five individual healthy volunteers on three separate days of the week with triplicate, back-to-back donations each day. Here, we report a total of 235 quantifiable proteins corresponding to 1,877 non-redundant peptides for evaluating sample collection reproducibility and establishing a healthy (human host) baseline EBC biomarker proteome studies. This work will pave the way for further investigations of EBC protein expression profiles and showcase the value of using non-invasive collection method techniques for clinically relevant biomarker discovery. This research was supported by the LDRD Biomedical Resilience And Readiness in AdVerse Operating Environments (BRAVE) Project (73748), and was conducted at Pacific Northwest National Laboratory (PNNL) in Richland, WA. PNNL is a multiprogram national laboratory operated by Battelle for the Department of Energy (DOE) under Contract DE-AC05-76RLO 1830.

59 BASIC BIOLOGICAL SCIENCES↗