Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data imbalance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

DOC-DICAM: Domain Aware One Class Defect Identification in Composite Aerostructure Material

Fiber-reinforced composites are a common material used in the design of aircraft structures due to their good tensile strength and resistance to compression. During the manufacturing process, these structures are thoroughly inspected for flaws and defects to ensure structural integrity during commercial use. Non-destructive testing (NDT) is a collection of inspection methods that allow inspectors to evaluate material without altering it. Due to the high safety standards in aerospace manufacturing, the NDT process is done manually and can be a significant bottleneck in the development workflow. In this paper, we develop an AI-based assistance tool to drastically reduce inspection time. Typical AI workflows require large amounts of annotated data, but defects rarely occur resulting in strong class imbalance. To overcome this, we formulate the problem of defect identification as an anomaly detection task in which our primary focus is learning non-defect characteristics. To do this, we develop a multi-task self-supervised learning framework that embeds problem specific domain knowledge into the deep learning model. We verify our method using fuselage data generated in a production environment. As a result, we show that our method can effectively identify defects and requires minimal training and inference time.

anomaly detection↗

Maven: a multimodal foundation model for supernova science

Abstract A common setting in astronomy is the availability of a small number of high-quality observations, and larger amounts of either lower-quality observations or synthetic data from simplified models. Time-domain astrophysics is a canonical example of this imbalance, with the number of supernovae observed photometrically outpacing the number observed spectroscopically by multiple orders of magnitude. At the same time, no data-driven models exist to understand these photometric and spectroscopic observables in a common context. Contrastive learning objectives, which have grown in popularity for aligning distinct data modalities in a shared embedding space, provide a potential solution to extract information from these modalities. We present Maven, the first foundation model for supernova science. To construct Maven, we first pre-train our model to align photometry and spectroscopy from 0.5 M synthetic supernovae using a contrastive objective. We then fine-tune the model on 4702 observed supernovae from the Zwicky transient facility. Maven reaches state-of-the-art performance on both classification and redshift estimation, despite the embeddings not being explicitly optimized for these tasks. Through ablation studies, we show that pre-training with synthetic data improves overall performance. In the upcoming era of the Vera C. Rubin observatory, Maven will serve as a valuable tool for leveraging large, unlabeled and multimodal time-domain datasets.

Zhang, Gemma (ORCID:0000000280198082)↗

Synthetic Infrasound Data for Machine Learning Detectors

Synthetic data is a powerful tool to generate large amounts of training data for machine learning models. The methods outlined in this report will be used to retrain the deep learning classifier for increased accuracy. Synthetic data will be useful to address the natural class imbalance between the different categories in the original ML work. Additionally, these tools will be applied for a variety of signal analysis methods that would use signals with a known signal-to-noise ratio for validation and testing.

58 GEOSCIENCES↗

The Persistent Challenge of Data Locality in the Post-Exascale Era

The era of exascale computing, exemplified by systems like Frontier achieving exaflop-level performance, marks a milestone. However, the quest for sheer compute power leads to strong imbalance in system design. Hence, scaling advancements in memory, network bandwidth, and storage are also necessary and pose challenges, with a crucial need to address data locality issues. This article underscores the fundamental importance of data locality as a key abstraction for optimizing application performance. Despite notable software solutions, the growing complexity of parallelism and memory hierarchy demands performance-portable data locality solutions across diverse computing platforms. Additionally, the article revisits data locality aspects, covering hardware considerations, application perspectives, software stack abstractions, and tool support. It concludes with insights into data locality challenges and opportunities, emphasizing the ongoing significance of collaborative research for progress in this critical issue.

Unat, Didem [Koc University, Istanbul (Turkey)] (O↗

Strangeness enhancement at its extremes: multiple (multi-)strange hadron production in pp collisions at \(\sqrt{s}=5.02\) TeV

The probability to observe a specific number of strange and multi-strange hadrons (nS), denoted as P(nS), is measured by ALICE at midrapidity (|y| < 0.5) in $$\sqrt{s}=5.02$$ TeV proton-proton (pp) collisions, dividing events into several multiplicity-density classes. Exploiting, for the first time, a technique based on counting the number of strange-particle candidates event-by-event, this measurement allows one to extend the study of strangeness production beyond the mean of the distribution. This constitutes a new test bench for production mechanisms, probing events with a large imbalance between strange and non-strange content. The analysis of a large-statistics data sample makes it possible to extract P(nS) up to a maximum nS of 7 for $${\text{K}}_{\text{S}}^{0}$$, 5 for Λ and $$\overline{\Lambda }$$, 4 for Ξ− and $${\overline{\Xi } }^{+}$$, and 2 for Ω− and $${\overline{\Omega } }^{+}$$. From this, the probability of producing strange hadron multiplets per event is calculated, thereby enabling the extension of the study of strangeness enhancement to extreme situations where several strange quarks hadronize in a single event at midrapidity. Moreover, comparing hadron combinations with different u and d quark compositions and equal overall s quark content, the contribution to the enhancement pattern coming from non-strangeness related mechanisms is isolated. The results are compared with state-of-the-art phenomenological models implemented in commonly used Monte Carlo event generators, including PYTHIA 8 Monash 2013, PYTHIA 8 with QCD-based Color Reconnection and Rope Hadronization (QCD-CR + Ropes), and EPOS LHC, which incorporates both partonic interactions and hydrodynamic evolution. These comparisons show that the new approach dramatically enhances the sensitivity to the different underlying physics mechanisms modeled by each generator.

Abualrob, I J↗

Boosted decision tree reweighting of simulated neutrino interactions for O ( 1 ) GeV neutrino cross-section measurements

This paper illustrates a generic method for multidimensional reweighting of O ( 1 ) GeV neutrino interaction Monte Carlo samples. The reweighting is based on a boosted decision tree algorithm trained on high-dimensional space in detector final-state observables. This enables one generator’s events to be reweighted so that its reconstructed particle content and kinematics distributions, as well as detector efficiency, match those of a target model. The approach establishes an efficient way to reuse legacy Monte Carlo data, avoiding regeneration. As an example, we test its use in a measurement of transverse kinematic imbalance of the μ - and proton in charged-current quasielastic like ν μ events from the MINERvA experiment.

Lin, Z. [Rochester U.] (ORCID:0009000188903698)↗

Investigating the Impact of Temporal and Directional Traffic Distribution on Crash Frequencies

Safety Performance Functions (SPFs) are mathematical models that establish relationships between the frequency of various crash types and site-specific characteristics, serving as essential tools for traffic safety analysis and roadway design. Traditional SPFs, however, often overlook the temporal fluctuations in traffic flow (such as peak-hour surges) and directional imbalances between opposing traffic streams. These traffic patterns can exacerbate congestion, disrupt driver behavior, and create unexpected conflict points, potentially leading to increased crash frequencies and more severe accidents. In light of this gap, this study aims to explore the potential of incorporating K-factors (representing peak-hour traffic proportions) and D-factors (reflecting the imbalance of directional traffic) into the development of SPFs to assess whether these factors can effectively represent the impact of temporal and spatial traffic distribution on roadway safety. Using crash data from Pennsylvania urban-suburban collector roadways, it is found that the D-factor plays a significant role in predicting the frequency of total crashes, fatal + injury crashes, and angle crashes, with positive coefficient signs indicating that higher directional imbalances correspond to increased crash risks. Similarly, the K-factor emerges as a critical predictor for fatal + injury crashes and rear-end crashes, with negative coefficients suggesting that a more pronounced traffic peak is associated with a reduction in expected crash frequencies. These results highlight the importance of accounting for uneven traffic distribution in both time and direction when developing SPFs, offering deeper insights into crash patterns and supporting more effective safety interventions and roadway designs.

Xu, Guanhao [ORNL] (ORCID:0000000214326357)↗

Critical statistical assessment of data in metal additive manufacturing

Obtaining high quality data reflecting the relationships between the additive manufacturing (AM) process parameters, material microstructure and mechanical properties is crucial for the use of machine learning in AM. A database of over 4,000 data entries of metal AM was created thanks to a large number of literature studies on key process parameters and indicators of build quality. Meta-analysis reveals critical biases in the literature. Firstly, majority of studies report only high quality builds, these imbalances in reporting result in weak correlation between process parameters, properties and consolidation, limiting the ability of machine learning models to generalize beyond optimized conditions. Nevertheless, the trained models accurately predict yield strength ($R^2 = 0.85$), suggesting that certain process–property relationships are effectively captured within these models. Secondly, quantitative microstructural data are largely absent, limiting the learning of the microstructure-mechanical properties relationships. Finally, current process window identification is based largely on the consolidation, despite significant uncertainty in its measurement. It is important to identify the process map on the basis of not only the consolidation, but also mechanical behaviour under loading. Such a identification shows that 316 L and Inconel have much larger process map (i.e. highly printable) in comparison to the AlSi10Mg and Ti6Al4V.

Additive manufacturing↗

Indicators of Global Climate Change 2023: annual update of key indicators of the state of the climate system and human influence

Intergovernmental Panel on Climate Change (IPCC) assessments are the trusted source of scientific evidence for climate negotiations taking place under the United Nations Framework Convention on Climate Change (UNFCCC). Evidence-based decision-making needs to be informed by up-to-date and timely information on key indicators of the state of the climate system and of the human influence on the global climate system. However, successive IPCC reports are published at intervals of 5–10 years, creating potential for an information gap between report cycles. We follow methods as close as possible to those used in the IPCC Sixth Assessment Report (AR6) Working Group One (WGI) report. We compile monitoring datasets to produce estimates for key climate indicators related to forcing of the climate system: emissions of greenhouse gases and short-lived climate forcers, greenhouse gas concentrations, radiative forcing, the Earth's energy imbalance, surface temperature changes, warming attributed to human activities, the remaining carbon budget, and estimates of global temperature extremes. The purpose of this effort, grounded in an open-data, open-science approach, is to make annually updated reliable global climate indicators available in the public domain. As they are traceable to IPCC report methods, they can be trusted by all parties involved in UNFCCC negotiations and help convey wider understanding of the latest knowledge of the climate system and its direction of travel. The indicators show that, for the 2014–2023 decade average, observed warming was 1.19 [1.06 to 1.30] °C, of which 1.19 [1.0 to 1.4] °C was human-induced. For the single-year average, human-induced warming reached 1.31 [1.1 to 1.7] °C in 2023 relative to 1850–1900. The best estimate is below the 2023-observed warming record of 1.43 [1.32 to 1.53] °C, indicating a substantial contribution of internal variability in the 2023 record. Human-induced warming has been increasing at a rate that is unprecedented in the instrumental record, reaching 0.26 [0.2–0.4] °C per decade over 2014–2023. This high rate of warming is caused by a combination of net greenhouse gas emissions being at a persistent high of 53±5.4 Gt CO 2 e yr -1 over the last decade, as well as reductions in the strength of aerosol cooling. Despite this, there is evidence that the rate of increase in CO 2 emissions over the last decade has slowed compared to the 2000s, and depending on societal choices, a continued series of these annual updates over the critical 2020s decade could track a change of direction for some of the indicators presented here.

54 ENVIRONMENTAL SCIENCES↗

Frequency-Nadir-Constrained Unit Commitment for Low-Inertia, High-IBR Island Power Systems [Slides]

The process of energy decarbonization in island power systems is accelerated due to the swift integration of inverter-based renewable energy resources (IBRs). The unique features of such systems, including rapid frequency changes resulting from potential generation outages or imbalances due to the unpredictability of renewable power, pose a significant challenge in maintaining the frequency nadir without external support. This paper presents a unit commitment (UC) model with data-driven frequency nadir constraints, including either frequency nadir or minimum inertia requirements, helping to limit frequency deviations after significant generator outages. The constraints are formulated using a linear regression model that takes advantage of real-world, year-long generation scheduling and dynamic simulation data. The efficacy of the proposed UC model is verified through a year-long simulation in an actual island power system using historical weather data. The alternative minimum inertia constraint, derived from actual system operation assumptions, is also evaluated. Findings demonstrate that the proposed frequency nadir constraint notably improves the system's frequency nadir under high photovoltaic (PV) penetration levels, albeit with a slight increase in generation costs, when compared to the alternative minimum inertia constraint.

14 SOLAR ENERGY↗

Search for dark matter produced in association with one or two top quarks in proton-proton collisions at $\sqrt{\text{s}}$ = 13 TeV

A search is performed for dark matter (DM) produced in association with a single top quark or a pair of top quarks using the data collected with the CMS detector at the LHC from proton-proton collisions at a center-of-mass energy of 13 TeV, corresponding to 138 fb −1 of integrated luminosity. An excess of events with a large imbalance of transverse momentum is searched for across 0, 1 and 2 lepton final states. Novel multivariate techniques are used to take advantage of the differences in kinematic properties between the two DM production mechanisms. No significant deviations with respect to the standard model predictions are observed. The results are interpreted considering a simplified model in which the mediator is either a scalar or pseudoscalar particle and couples to top quarks and to DM fermions. Axion-like particles that are coupled to top quarks and DM fermions are also considered. Expected exclusion limits of 410 and 380 GeV for scalar and pseudoscalar mediator masses, respectively, are set at the 95% confidence level. A DM particle mass of 1 GeV is assumed, with mediator couplings to fermions and DM particles set to unity. A small signal-like excess is observed in data, with the largest local significance observed to be 1.9 standard deviations for the 150 GeV pseudoscalar mediator hypothesis. Because of this excess, mediator masses are only excluded below 310 (320) GeV for the scalar (pseudoscalar) mediator. The results are also translated into model-independent 95% confidence level upper limits on the visible cross section of DM production in association with top quarks, ranging from 1 pb to 0.02 pb.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Base excision repair and double strand break repair cooperate to modulate the formation of unrepaired double strand breaks in mouse brain

Abstract We lack the fundamental information needed to understand how DNA damage in the brain is generated and how it is controlled over a lifetime in the absence of replication check points. To address these questions, here, we integrate cell-type and region-specific features of DNA repair activity in the normal brain. The brain has the same repair proteins as other tissues, but normal, canonical repair activity is unequal and is characterized by high base excision repair (BER) and low double strand break repair (DSBR). The natural imbalance creates conditions where single strand breaks (SSBs) can convert to double strand breaks (DSBs) and reversibly switch between states in response to oxidation both in vivo and in vitro. Our data suggest that, in a normal background of repair, SSBs and DSBs are in an equilibrium which is pushed or pulled by metabolic state. Interconversion of SSB to DSBs provides a physiological check point, which would allow the formation of unrepaired DSBs for productive functions, but would also restrict them from exceeding tolerable limits.

Science & Technology - Other Topics↗

Measurement of muon neutrino charged-current quasielasticlike cross section using off-axis NuMI beam at ICARUS

This paper presents the first neutrino cross-section measurement from the ICARUS detector at Fermilab, using Neutrinos at the Main Injector (NuMI) beam data collected from two beam operation periods corresponding to 2.5 × 10 20 protons-on-target in neutrino beam mode. The signal is defined by events with no pions produced in the final state, a topology dominated by charged-current quasielasticlike signatures. The measurement is reported as flux-averaged differential cross sections as functions of kinematic variables that provide sensitivity to the complex nuclear effects that often dominate the systematic uncertainty budgets of neutrino oscillation measurements. Specifically, this work reports cross sections in two angular variables—the angle of the outgoing lepton and the opening angle between the lepton and leading proton—and two variables characterizing the kinematic imbalance between the muon and proton in the plane transverse to the incoming neutrino. These results are compared against predictions from a variety of neutrino event generators, with p values calculated between the extracted cross sections and each prediction. Overall, the predictions agree with the data; however, the current budget of uncertainties does not yet provide sufficient discriminating power to favor a specific model.

Abd Alrahman, F. [Houston U.]↗

Measurement-Based Approach for Inertia-Trend Analysis of the US Western Interconnection

Rising deployment of inverter-based resources (IBRs), characterized by a lack of rotating mass, is decreasing the total inertia of the system. This can lead to an increased Rate of Change of Frequency (RoCoF) during the disturbance and false activation of protective devices. There is a need to assess the inertia over the past decade amidst the evolving landscape of renewable energy sources to develop strategies for integrating energy storage, enhancing resilience measures, and ensuring the stable and reliable operation of the grid. Therefore, a realistic assessment of the inertia trend using a measurement-based approach that addresses the limitations of existing models is proposed. An inertia study of the Western Interconnection in the United States is performed utilizing the data from 2013 to 2022, obtained from FNET/ GridEye network. The three-second RoCoF time window is chosen for the study as it showed an optimum balance between a strong correlation with the power imbalance (ΔP) and minimum inclusion of primary response from governor. The obtained inertia trend result shows a small percentage declination of inertia over the decade. By examining the result alongside a generation mix graph, insights are gained into the dynamic interplay between shifting energy landscape and system inertia.

Dulal, Saurav↗

Measurement of the A dependence of the ν μ charged-current quasielasticlike cross section as a function of muon and proton kinematics at ⟨E ν ⟩ ∼ 6 GeV

The first simultaneous measurements of the 𝜈 𝜇 quasielasticlike cross section on C, CH, H 2 ⁡O, Fe, and Pb targets as a function of kinematic imbalance variables in the plane transverse to the incoming neutrino direction are presented. These variables combine the muon and proton information to provide a new way to disentangle the effects of the nucleus in quasielasticlike processes. The data were obtained using a wideband 𝜈 𝜇 beam with ⟨E 𝜈 ⟩ ∼ 6 GeV. Cross-section ratios of the different target materials to CH are also shown. These measurements are used to explore the nature of the cross-section 𝐴 scaling, as well as initial and final state interaction effects. Comparisons are made to predictions from a number of commonly used neutrino Monte Carlo event generators. The range of predictions of the different models tends to cover the data but the degree and consistency of the agreement suffers in regions, and on higher 𝐴 targets, where the final state interactions are expected to be more pronounced.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

Topography-Induced TKE Budget Behavior Over an Amazon Forest

Data from three different heights (35, 50 and 81 m) of one of the ATTO Project towers in the Amazon forest were used to calculate the TKE (turbulence kinetic energy) budget and some other statistics within the RSL (roughness sublayer). The statistical analyses were carried out for unstable and stable cases. The vertical transport and the vertical advection terms do not explain the imbalances found in the TKE budget, highlighting the impor- tance of horizontal terms in complex terrain. Indeed, patterns independent of the time of day were found for vertical transport, horizontal and vertical velocity skewness, the mean vertical component velocity, and momentum and vertical TKE flux. Here, the highest values of mean vertical velocity were observed from the direction where there is a valley near the tower, which may indicate a topographical effect. Monin-Obukhov Similarity Theory (MOST) predictions for the mean velocity gradient, strictly applicable only in the inertial sublayer (ISL), gave reasonable predictions at 35 and 50 m, but dimensionless dissipation rates displayed large scatter. At 81 m, there were clear deviations from MOST for the two functions, disclosing the effect of topography. In addition, a neutral LES (large-eddy simu- lation) study showed good agreement with tower data for the same general wind direction.

ATTO project↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗