Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Human Error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Evaluating algorithmic bias on biomarker classification of breast cancer pathology reports

Objectives: This work evaluated algorithmic bias in biomarkers classification using electronic pathology reports from female breast cancer cases. Bias was assessed across 5 subgroups: cancer registry, race, Hispanic ethnicity, age at diagnosis, and socioeconomic status. Materials and Methods: We utilized 594 875 electronic pathology reports from 178 121 tumors diagnosed in Kentucky, Louisiana, New Jersey, New Mexico, Seattle, and Utah to train 2 deep-learning algorithms to classify breast cancer patients using their biomarkers test results. We used balanced error rate (BER), demographic parity (DP), equalized odds (EOD), and equal opportunity (EOP) to assess bias. Results: We found differences in predictive accuracy between registries, with the highest accuracy in the registry that contributed the most data (Seattle Registry, BER ratios for all registries >1.25). BER showed no significant algorithmic bias in extracting biomarkers (estrogen receptor, progesterone receptor, human epidermal growth factor receptor 2) for race, Hispanic ethnicity, age at diagnosis, or socioeconomic subgroups (BER ratio <1.25). DP, EOD, and EOP all showed insignificant results. Discussion: We observed significant differences in BER by registry, but no significant bias using the DP, EOD, and EOP metrics for socio-demographic or racial categories. This highlights the importance of employing a diverse set of metrics for a comprehensive evaluation of model fairness. Conclusion: A thorough evaluation of algorithmic biases that may affect equality in clinical care is a critical step before deploying algorithms in the real world. We found little evidence of algorithmic bias in our biomarker classification tool. Artificial intelligence tools to expedite information extraction from clinical records could accelerate clinical trial matching and improve care.

60 APPLIED LIFE SCIENCES↗

Issue Resolution During the Development of the Performance Assessment for the Savannah River Site Saltstone Disposal Facility - 20125

In 2019, Savannah River Remediation developed a revision to the performance assessment (PA) on behalf of the U.S. Department of Energy (DOE) Savannah River Operations Office (SR) for the near-surface disposal of low-level waste at the Savannah River Site (SRS) Saltstone Disposal Facility (SDF). Soluble waste from SRS Tank Farms undergoes salt processing to remove cesium and other high-activity constituents. The low-activity decontaminated salt solution (DSS) is then immobilized by mixing it into a cementitious waste form known as saltstone. After mixing, the saltstone is poured into leak-tight concrete vaults, known as saltstone disposal units (SDUs), where the waste form cures. By the time of facility closure, the SDF is expected to consist of 15 SDUs with a combined capacity of 1.06 E+09 L (280 Mgal) of cured saltstone. The facility operates under a Disposal Authorization Statement from DOE and a permit from the South Carolina Department of Health and Environmental Control (SCDHEC). Since the start of operations in 1990, the SDF has received almost 6.7 E+07 L (18 Mgal) of DSS, resulting in the safe disposal of 2.7 E+16 Bq (7.3 E+05 Ci) of activity. Due to the radioactive decay of short-lived contaminants, the total remaining activity in the disposed waste is estimated to be approximately 1.4 E+16 Bq (3.9 E+05 Ci), as of September 2018. The Disposal Authorization Statement requires a demonstration that the system of engineered and natural features of the disposal facility will limit releases from the facility and be protective of human health and the environment for at least the next 1,000 years. The long-term performance of the facility was evaluated under the requirements of the DoE's Radioactive Waste Management Manual (US DOE Manual 435.1-1). Simulations were performed to demonstrate that the disposal facility would meet performance objectives specified in the manual. The evaluation was based on numerical models that simulate the releases of contaminants from the saltstone waste form. Contaminants were transported through groundwater and air pathways to points of assessment to evaluate compliance (i.e., 100 m from the SDUs). In addition, the potential consequences of an inadvertent human intrusion (IHI) were also evaluated. A number of issues were overcome during the development of these simulations. These issues were identified as part of internal technical reviews. Simulations are developed by people and people make mistakes, so the internal technical review process is a vital step in PA development. Specific examples of resolved issues include a unit-conversion error, an inappropriate definition for a model boundary condition, a model time-stepping issue, and an error in the calculation for the buildup of contaminants in soil. Actions taken to address these issues resulted in an improved product with a better supported technical basis and more defensible results. The identification and correction of these issues are discussed. By understanding these issues, model developers and technical reviewers working on PAs in the future may avoid repeating these types of mistakes. Transparency with respect to these mistakes builds trust between waste management sites, regulators, and stakeholders. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A Combined Computer Vision and Deep Learning Approach for Rapid Drone-Based Optical Characterization of Parabolic Troughs

Optical accuracy is a primary driver of parabolic trough concentrating solar power (CSP) plant performance, but can be damaged by wind loads, gravity, error during installation, and regular plant operation. Collecting and analyzing optical measurements over an entire operating parabolic trough plant is difficult, given the large scale of typical installations. Distant Observer, a software tool developed at the National Renewable Energy Laboratory, uses images of the absorber tube reflected in the collector mirror to measure both surface slope in the parabolic mirror and offset of the absorber tube from the ideal focal point. This technology has been adapted for fast data collection using low-cost commercial drones, but until recently still required substantial human labor to process large amounts of data. A new method leveraging advanced deep learning and computer vision tools can drastically reduce the time required to process images. This new method addresses the primary analysis bottleneck, identifying featureless, reflective mirror corner points to a high degree of accuracy. Recent work has shown promising results using computer vision methods. The combined deep learning and computer vision approach presented here proved highly effective and has the potential to further automate data collection and analysis, making the tool more robust. The method presented in this paper automatically identified 74.3% of mirror corners within 2 pixels of their manually marked counterparts and 91.9% within 3 pixels. This level of accuracy is sufficient for practical Distant Observer analysis within a target uncertainty. A commercial drone collected video of over 100 parabolic trough modules at an operating CSP plant to demonstrate the deep learning and computer vision method's usefulness in processing large amounts of data. These troughs were successfully analyzed using Distant Observer, paired with the new deep learning and computer vision algorithm, and can provide plant operators and trough designers with valuable insight about plant performance, operating strategies, and plant-wide optical error trends.

computer vision↗

Performance-based earthquake early warning for tall buildings

The ShakeAlert Earthquake Early Warning (EEW) system aims to issue an advance warning to residents on the West Coast of the United States seconds before the ground shaking arrives, if the expected ground shaking exceeds a certain threshold. However, residents in tall buildings may experience much greater motion due to the dynamic response of the buildings. Therefore, there is an ongoing effort to extend ShakeAlert to include the contribution of building response to provide a more accurate estimation of the expected shaking intensity for tall buildings. Currently, the supposedly ideal solution of analyzing detailed finite element models of buildings under predicted ground-motion time histories is not theoretically or practically feasible. The authors have recently investigated existing simple methods to estimate peak floor acceleration (PFA) and determined these simple formulas are not practically suitable. Instead, this article explores another approach by extending the Pacific Earthquake Engineering Research Center (PEER) performance-based earthquake engineering (PBEE) to EEW, considering that every component involved in building response prediction is uncertain in the EEW scenario. Additionally, while this idea is not new and has been proposed by other researchers, it has two shortcomings: (1) the simple beam model used for response prediction is prone to modeling uncertainty, which has not been quantified, and (2) the ground motions used for probabilistic demand models are not suitable for EEW applications. In this article, we address these two issues by incorporating modeling errors into the parameters of the beam model and using a new set of ground motions, respectively. We demonstrate how this approach could practically work using data from a 52-story building in downtown Los Angeles. Using the criteria and thresholds employed by previous researchers, we show that if peak ground acceleration (PGA) is accurately estimated, this approach can predict the expected level of human comfort in tall buildings.

58 GEOSCIENCES↗

Development of a Convolutional Neural Network Classifier for Data Starved Spectra - 20199

The Institute for Clean Energy Technology (ICET) at Mississippi State University is exploring the utility of machine learning in augmenting its mobile radiation surveying platforms, which are currently being developed as means to survey depleted uranium contaminated areas in support of remediation and decommissioning efforts. Mobile survey platforms provide a means to efficiently scan large areas of interest while reducing human exposure to radiation and other hazards. The survey platforms can also be used for scanning for any gamma emitting isotope in addition to depleted uranium. The spectral data that the platforms collect may be data starved with relatively low counts and poorly defined spectral features depending on the speed of the platforms and scintillation detector selection. Such data-starved spectra are difficult to use for isotope identification, requiring advanced knowledge of the possible radionuclides that could be present and environmental factors that could attenuate signals or introduce background noise. These factors in combination with the volume of survey data increases the time it takes to perform analysis of survey data when the source type is unknown. There are a number of algorithms in the field of machine learning that can be used to classify data that would be challenging and time-consuming for a human to identify. Supervised machine learning algorithms train models based on extensive amounts of human-labeled training data. Once sufficiently trained, these models can be used to quickly make high-fidelity predictions on new data. Convolutional neural networks are machine learning algorithms that excel in learning representations of 'shapes'. They do this by taking numerical input data and convolving them with spatial feature detectors referred to as filters. These filters are incrementally adjusted to reduce the prediction error on the data during the backpropagation step of training. Discussed in this paper is the development of a convolutional neural network classifier (CNNC) that can utilize spectral survey data for source discrimination and isotope identification. Bench-top laboratory experiments data using LaBr{sub 3}(Ce) scintillation detectors were used to train and evaluate the performance of the developed CNNC. The CNNC is capable of discriminating a variety of gamma emitting source types, differentiating different forms of uranium (depleted vs. natural), and estimating the amount of uranium for a known geometry. The discussed CNNC may be useful in scenarios where survey systems are deployed in situations where hazardous radioactive material maybe present, but the type is unknown. When used in remediation applications the CNNC can be used to screen-out false positives, helping reduce remediation costs. (authors)

07 ISOTOPE AND RADIATION SOURCES↗

Data Collection and Analysis Challenges and Mitigation Strategies for Quantitative Human Factors Research Studies in Nuclear Power Plant Modernization

The United States (U.S.) Department of Energy (DOE) Light Water Reactor Sustainability (LWRS) program Plant Modernization Pathway is conducting targeted research and development (R&D) to address aging and reliability concerns with the legacy instrumentation and control and related information systems of the U.S. LWR fleet. In this effort, the application of human factors engineering (HFE) provides an important role in ensuring new digital plant technologies enable broad innovation and business improvement with continued operational safety. Evaluation is a key activity in HFE, which often occurs iteratively through the system design lifecycle. While qualitative methods are important in collecting information of how users perform tasks through observations, quantitative methods are equally important in assessing system design based on performance. In collecting and analyzing this quantitative data, there are notable challenges in the nuclear HFE domain that may threaten the validity and reliability of the inferences made in these studies. Notable challenges include small sample size and limited resources, large error variance and small effect size, an “adding test features to losing degrees of freedom” dilemma, non-normal distribution, and heterogeneity of variance. The results in control room usability studies are often statistically non-significant, which makes it hard to interpret. This work discusses these challenges across different scientific viewpoints and provides real-world examples of these challenges in practice. Collectively, the objective of this work is to position these challenges to the larger data science community as a means of identifying future opportunities to address these issues.

99 GENERAL AND MISCELLANEOUS↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

I Know I'm Right, But Does My Phone?

Transportation is the largest source of green-house gas emissions in the United States. Reducing transportation emissions depends on human travel behavior, which relies on local land use and planning. Travel diaries, consisting of sequences of trips between places for a particular individual, are typically used to instrument human travel behavior. However, these diaries are only as accurate as the underlying methods used to construct them. Travel diary algorithms have been a popular research topic since the advent of GPS tracking surveys. Mode inference algorithms in particular have been well represented in literature. However, these algorithms have typically been validated using prompted recall of pre-segmented trips, which doesn't account for segmentation error, thus disregarding the continuity of mode inference. Furthermore, phone operating systems and applications have adopted battery-conserving techniques, but we are not aware of prior work that has characterized the resulting data collection errors or evaluated procedures to mitigate them. We introduce a framework to evaluate accuracy of trip length computations and mode inference. We develop a temporal alignment procedure in analyzing continuous mode-segmented trajectories for groups of trips. We then apply our framework to evaluate an example set of travel diary algorithms from the open-source OpenPATH travel diary platform against MobilityNet, a public dataset containing information from three artificial timelines that cover 15 different travel modes. Our results show that inference based on an integration with map features results in weighted F_1 scores of 0.60 (iOS) and 0.74 (android). We also show that OpenPATH tends to under count trip length, with mean of signed relative error of -0.0438 on android and -0.0704 on iOS. We hope that other travel diary algorithms will be evaluated using this standardized process, and that the results used to understand and improve the state-of-the-art in this field.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Advocating Feedback Control for Human-Earth System Applications

This paper proposes a feedback control perspective for Human-Earth Systems (HESs) which essentially are complex systems that capture the interactions between humans and nature. Recent attention in HES research has been directed towards devising strategies for climate change mitigation and adaptation, aimed at achieving environmental and societal objectives. However, existing approaches heavily rely on HES models, which inherently suffer from inaccuracies due to the complexity of the system. Moreover, overly detailed models often prove impractical for optimization tasks. We propose a framework inheriting from feedback control strategies the robustness against model errors, because inaccuracies are mitigated using measurements retrieved from the field. The framework comprises two nested control loops. The outer loop computes the optimal inputs to the HES, which are then implemented by actuators controlled in the inner loop. Potential fields of applications are also identified and a numerical example is provided.

biological system modeling↗

CrossCheck: Rapid, Reproducible, and Interpretable Model Evaluation

Evaluation beyond aggregate performance metrics, e.g. F1-score, is crucial to both establish an appropriate level of trust in machine learning models and identify future model improvements. In this paper we demonstrate CrossCheck, an interactive visualization tool for rapid crossmodel comparison and reproducible error analysis. We describe the tool and discuss design and implementation details. We then present three use cases (named entity recognition, reading comprehension, and clickbait detection) that show the benefits of using the tool for model evaluation. CrossCheck allows data scientists to make informed decisions to choose between multiple models, identify when the models are correct and for which examples, investigate whether the models are making the same mistakes as humans, evaluate models’ generalizability and highlight models’ limitations, strengths and weaknesses. Furthermore, CrossCheck is implemented as a Jupyter widget, which allows rapid and convenient integration into data scientists’ model development workflows.

Arendt, Dustin L.↗

Using machine learning with optical profilometry for GaN wafer screening

Abstract To improve the manufacturing process of GaN wafers, inexpensive wafer screening techniques are required to both provide feedback to the manufacturing process and prevent fabrication on low quality or defective wafers, thus reducing costs resulting from wasted processing effort. Many of the wafer scale characterization techniques—including optical profilometry—produce difficult to interpret results, while models using classical programming techniques require laborious translation of the human-generated data interpretation methodology. Alternatively, machine learning techniques are effective at producing such models if sufficient data is available. For this research project, we fabricated over 6000 vertical PiN GaN diodes across 10 wafers. Using low resolution wafer scale optical profilometry data taken before fabrication, we successfully trained four different machine learning models. All models predict device pass and fail with 70–75% accuracy, and the wafer yield can be predicted within 15% error on the majority of wafers.

36 MATERIALS SCIENCE↗

Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning

Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.

54 ENVIRONMENTAL SCIENCES↗

Modeling of streamflow in a 30 km long reach spanning 5 years using OpenFOAM 5.x

Abstract. Developing accurate and efficient modeling techniques for streamflow at the tens-of-kilometers spatial scale and multi-year temporal scale is critical for evaluating and predicting the impact of climate- and human-induced discharge variations on river hydrodynamics. However, achieving such a goal is challenging because of limited surveys of streambed hydraulic roughness, uncertain boundary condition specifications, and high computational costs. We demonstrate that accurate and efficient three-dimensional (3-D) hydrodynamic modeling of natural rivers at 30 km and 5-year scales is feasible using the following three techniques within OpenFOAM, an open-source computational fluid dynamics platform: (1) generating a distributed hydraulic roughness field for the streambed by integrating water-stage observation data, a rough wall theory, and a local roughness optimization and adjustment strategy; (2) prescribing the boundary condition for the inflow and outflow by integrating precomputed results of a one-dimensional (1-D) hydraulic model with the 3-D model; and (3) reducing computational time using multiple parallel runs constrained by 1-D inflow and outflow boundary conditions. Streamflow modeling for a 30 km long reach in the Columbia River (CR) over 58 months can be achieved in less than 6 d using 1.1 million CPU hours. The mean error between the modeled and the observed water stages for our simulated CR reach ranges from −16 to 9 cm (equivalent to approximately ±7 % relative to the average water depth) at seven locations during most of the years between 2011 and 2019. We can reproduce the velocity distribution measured by the acoustic Doppler current profiler (ADCP). The correlation coefficients of the depth-averaged velocity between the model and ADCP measurements are in the range between 0.71 and 0.83 at 75 % of the survey cross sections. With the validated model, we further show that the relative importance of dynamic pressure versus hydrostatic pressure varies with discharge variations and topography heterogeneity. Given the model's high accuracy and computational efficiency, the model framework provides a generic approach to evaluate and predict the impacts of climate- and human-induced discharge variations on river hydrodynamics at tens-of-kilometers and decadal scales.

58 GEOSCIENCES↗

Automated vehicle microscopic energy consumption study (AV-Micro): Data collection and model development

While the Adaptive Cruise Control (ACC) system in automated vehicles (AVs) is expected to impact transportation energy significantly, existing AV energy consumption models only directly adopt those developed with Human-driven Vehicle (HV) data without even slight adaptation or calibration to accommodate unique AV energy consumption features. This study will investigate how accurately HV data-based models can predict the energy consumption of AVs. Empirical trajectory data and corresponding instantaneous energy consumption rates from both AVs and HVs were collected. We adopted two classical HV data-based models to fit these data. The calibration results indicated that these models yield around 20 30% prediction errors for AVs. To further improve the prediction accuracy, this study designed an AV-Micro model by incorporating components of multiple classic energy consumption models that better capture ACC energy consumption features, including piecewise driving behavior. With this, the AV-Micro model achieves lower than 10% prediction errors. The AV-Micro model’s high consistency across different test runs was verified with statistical significance tests, demonstrating its adaptability in different driving profiles. To confirm the discrepancies between the energy consumption features of AVs and HVs, more statistical significance tests were conducted to show that the AV-Micro model cannot be directly applied to HV data. The findings by calibrated AV-Micro models revealed that AVs consume approximately 80.5–146.4 J more energy than HVs for each meter traveled. Furthermore, the frequency analysis of energy consumption indicates that there is still some room for AVs to improve energy efficiency, particularly given their larger amplitude high-frequency fluctuations.

33 ADVANCED PROPULSION SYSTEMS↗

Accurate Identification of Deamidation and Citrullination from Global Shotgun Proteomics Data Using a Dual-Search Delta Score Strategy

While proteins with deamidated/citrullinated amino acids play critical roles in the pathogenesis of many human diseases, identifying these modifications in complex biological samples has been an ongoing challenge. Herein we present a method to accurately identify these modifications from shotgun proteomics data from a deep proteome profiling study of human pancreatic islets obtained by laser capture microdissection. All MS/MS spectra were searched against database by MSGF+ twice with or without a +0.9840 Da mass shift on amino acids asparagine, glutamine, and arginine (NQR) as a dynamic modification. Consequently, for each spectrum the resulting two peptide-to-spectrum matches (PSM) with their respective MSGF+ scores were used for Delta Score calculation. It was observed that all PSMs with positive Delta Score values were clustered with mass errors around 0 ppm, while PSMs with negative Delta Score values were distributed nearly equally within the defined mass error range (20 ppm) for database searching. To estimate false discovery rate (FDR), a “pseudo-decoy” approach was applied whether datasets were searched against a database with a “real modification” mass shift (+0.9840 Da) and a “mock modification” mass shift (+1.0227 Da). FDR was controlled to ~2% with a Delta Score filter greater than zero. Manual inspection of spectra showed that PSMs with positive Delta Score value contained deamidated/citrullinated fragments in their MS/MS spectra. Finally, the results demonstrated that in-situ deamidated/citrullinated peptides can be accurately identified from shotgun tissue proteomics data by the Dual-search Delta Score Strategy.

59 BASIC BIOLOGICAL SCIENCES↗

Measurement Uncertainty in One-of-a-kind Experiments

A golden standard in science is to repeat an experimental measurement multiple times and calculate the measured value with experimental uncertainty by following well developed statistical procedures. For various reasons – cost, technical difficulty, international treaties, ethics of dealing with human or animal subjects, ecology - many important experiments and observations can not be repeated. Astronomy, earthquakes, hurricanes produce data from one-of-a-kind events. When information is not available in any other way, it should not be dismissed as qualitative, anecdotal evidence only. Analyzed in a mathematically rigorous way, it produces quantitative experimental data. Analysis of data from one-of-a-kind event differs from analysis of repeated experiments’ data. For repeated experiments, the experimental error includes a range of true values generated by repetitions of the experiment, and measurement uncertainty caused by detectors. They are independent. Repetitions of any experiment, as similar as achievable, always have built-in differences resulting in a range of the true values rather than in a single value. Measurement uncertainty depends on the measurement system only. Digital measurements have very small uncertainty, frequently smaller than the range of true experimental values resulting from built-in differences in the experiment repetitions. When data from one–of–a kind experiment are analyzed, only the measurement uncertainty can be reported.

42 ENGINEERING↗