Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “MATHEMATICAL STATISTICS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Diffusion Codes: Self-Correction from Small(er)-Set Expansion with Tunable Non-locality

Optimal constructions of classical LDPC codes can be obtained by choosing the Tanner graph uniformly at random among biregular graphs. We introduce a class of codes that we call ``diffusion codes'', defined by placing each edge connecting bits and checks on some graph, and acting on that graph with a random SWAP network. By tuning the depth of the SWAP network, we can tune a tradeoff between the amount of randomness -- and hence the optimality of code parameters -- and locality with respect to the underlying graph. For diffusion codes defined on the cycle graph, if the SWAP network has depth $\sim Tn$ with $T> n^{2β}$ for arbitrary $β>0$, then we prove that almost surely the Tanner graph is a lossless ``smaller set'' vertex expander for small sets up size $δ\sim \sqrt T \sim n^β$, with bounded bit and check degree. At the same time, the geometric size of the largest stabilizer is bounded by $\sqrt T$ in graph distance. We argue, based on physical intuition, that this result should hold more generally on arbitrary graphs. By taking hypergraph products of these classical codes we obtain quantum LDPC codes defined on the torus with smaller-set boundary and co-boundary expansion and the same expansion/locality tradeoffs as for the classical codes. These codes are self-correcting and admit single-shot decoding, while having the geometric size of the stabilizer growing as an arbitrarily small power law. Our proof technique establishes mixing of a random SWAP network on small subsystems at times scaling with only the subsystem size, which may be of independent interest.

Combinatorics (math.CO)↗

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES↗

ASME Design Code Rule Changes for Nuclear Graphite

The American Society of Mechanical Engineers Boiler Pressure and Vessel Code (ASME BPVC) Section III, Division 5, Article HHA-3000 outlines graphite core component and graphite core assembly design guidelines. Graphite core components are defined as ?components manufactured from graphite that are installed to form a graphite core assembly within the reactor pressure vessel of a high temperature, graphite moderated fission reactor.? (p. 413) Graphites? inherent defect distributions do not allow for deterministic material reliability. Rather, graphite has variable strength distributions which change by grade. Article HHA-3000 outlines two semi-probabilistic methods, the full and simplified assessments, which set design load limit targets for each of three component structural reliability classes. The Design Task Group was officially recognized as a specialized task group within ASME November of 2023, though we?ve been collaborating since 2022. The purpose of the Design Task Group is to correct, clarify, and make HHA-3000 function as intended. The Design Task Group will sunset once we?ve achieved our objectives. The Design Task Group was specifically told to not write new Code. While there may be more precise and more accurate methods to determine reliability targets, the current methods are conservative, relatively simple to implement, and have thus far been considered satisfactory for setting design reliability targets. Much of the ground-work to write proposal files and background documents for records to make the changes needed to achieve our objective have been completed. The Design Task Group has documented much of their work through papers, presentations, and memorandums. Three memorandums in which INL team members had substantial contributions are found in the Appendices: FEA Modeling for the Baseline Program, Evaluating the Effects on Margin of Updating the Threshold and Shape Parameters in the Full Assessment, and Interpretations of the Full and Simplified Assessments in ASME BPVC. Most of the on-going work to achieve the Design Task Group?s objective will be addressing comments on existing records and moving records through the balloting process. The Design Task Group met bi-weekly mostly through the end of FY2023. Since February 2024, the Design Task Group has mostly been completed with solving and documenting the technical issues associated with the assessments. Unless new tasks are identified, the remaining work of the Design Task Group will be political and editorial.

97 MATHEMATICS AND COMPUTING↗

Statistical Estimation of EV Driver Charging Behavior and Influential Factors

INL received data collected via telematics from battery electric vehicles (BEVs), and these vehicles were owned by retail customers who had entered into a telematics user agreement. The goal of analyzing these data was to develop mathematical models to characterize how different sets of BEV drivers use charging infrastructure at home and away from home (i.e., public charging) and quantify how various factors influence BEV drivers’ decision to charge and use available infrastructure. The data used in this analysis are unique because they provide real world BEV driving and charging behavior at the individual driving and parking event level. In this study we seek to leverage this data to quantify BEV charging and driving metrics to help inform models that predict quantities like the specific times when loads are imposed on the electrical grid due to BEV charging. Most models that have been developed to predict electrical grid load due to BEV charging, use simulations of BEV driving events and rely on assumptions such as every vehicle charges every night. Using a statistical modelling framework, we seek to investigate BEV charging behavior and quantitatively assess these common assumptions of BEV charging behavior.

33 - ADVANCED PROPULSION SYSTEMS↗

Leveraging interpolation models and error bounds for verifiable scientific machine learning

Effective verification and validation techniques for modern scientific machine learning workflows are challenging to devise. Statistical methods are abundant and easily deployed, but often rely on speculative assumptions about the data and methods involved. Error bounds for classical interpolation techniques can provide mathematically rigorous estimates of accuracy, but often are difficult or impractical to determine computationally. Here, in this work, we present a best-of-both-worlds approach to verifiable scientific machine learning by demonstrating that (1) multiple standard interpolation techniques have informative error bounds that can be computed or estimated efficiently; (2) comparative performance among distinct interpolants can aid in validation goals; (3) deploying interpolation methods on latent spaces generated by deep learning techniques enables some interpretability for black-box models. We present a detailed case study of our approach for predicting lift-drag ratios from airfoil images. Code developed for this work is available in a public Github repository.

97 MATHEMATICS AND COMPUTING↗

Impacts of PV Module Connector Failures on Cost and Performance of Utility Scale Photovoltaic Systems

The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.

14 SOLAR ENERGY↗

Clustering and Cliques in Preferential Attachment Random Graphs with Edge Insertion

In this paper, we investigate the global clustering coefficient (a.k.a transitivity) and clique number of graphs generated by a preferential attachment random graph model with an additional feature of allowing edge connections between existing vertices. Specifically, at each time step t, either a new vertex is added with probability f(t), or an edge is added between two existing vertices with probability 1 – f(t). We establish concentration inequalities for the global clustering and clique number of the resulting graphs under the assumption that f(t) is a regularly varying function at infinity with index of regular variation –$\gamma$, where $\gamma$ $\in$ [0, 1). Finally, we also demonstrate an inverse relation between these two statistics: the clique number is essentially the reciprocal of the global clustering coefficient.

97 MATHEMATICS AND COMPUTING↗

Open World Dempster-Shafer Theory/The Transferable Belief Model with Intervals: A Practitioner's Guide to DST and TBM

Dempster-Shafer theory (DST) is a mathematical framework that allows for uncertainty or ignorance to be quantified and included when making predictions from evidence. This is in contrast to Bayesian theory, which does not allow for any quantification of ignorance. The framework is described in great detail in [7]. DST is particularly useful for problems where the inclusion of additional evidence (for example, data from another sensor) could lead to a different conclusion. Thus, it is a useful data fusion method, especially in applications not suited to maximum likelihood or maximum a posteriori estimations due to limited samples or incomplete prior knowledge.

97 MATHEMATICS AND COMPUTING↗

From Ensemble Climate to Ensemble Impacts

Many climate-risk tools rely on ensemble mean projections or endpoint climate snapshots to characterize future hazards. Although convenient for communication, these representations remove the statistical, temporal, and physical information that real infrastructure systems respond to. Infrastructure degradation and failure arise from extremes, sequences, cumulative stress, compound hazards, and nonlinear fragility relationships, none of which survive ensemble averaging or temporal compression. Power-system failure statistics and cascading failure models further show that infrastructure risk is dominated by tail events and path-dependent dynamics rather than by mean conditions. This paper demonstrates why ensemble mean or endpoint-only climate representations are mathematically and physically inconsistent with engineering-grade risk analysis. We outline a model-resolved, time-series-based workflow that preserves extremes, variability, and sequencing by propagating each climate-model realization independently through hazard formation, exposure, fragility, and cascading failure mechanisms. Taking the ensemble of impacts—rather than the ensemble of climate—provides a defensible, physically coherent foundation for infrastructure resilience planning, regulatory compliance, and long-term investment decisions.

54 - ENVIRONMENTAL SCIENCES/GLOBAL CLIMATE CHANGE ↗

Constraining the phase shift of relativistic species in DESI BAOs

In the early Universe, neutrinos decouple quickly from the primordial plasma and propagate without further interactions. The impact of free-streaming neutrinos is to create a temporal shift in the gravitational potential that impacts the acoustic waves known as baryon acoustic oscillations (BAOs), resulting in a non-linear spatial shift in the Fourier-space BAO signal. In this work, we make use of and extend upon an existing methodology to measure the phase shift amplitude $\beta _{\phi }$ and apply it to the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) BAOs with an anisotropic BAO fitting pipeline. We validate the fitting methodology by testing the pipeline with two publicly available fitting codes applied to highly precise cubic box simulations and realistic simulations representative of the DESI DR1 data. We find further study towards the methods used in fitting the BAO signal will be necessary to ensure accurate constraints on $\beta _{\phi }$ in future DESI data releases. Using DESI DR1, we present individual measurements of the anisotropic BAO distortion parameters and the $\beta _{\phi }$ for the different tracers, and additionally a combined fit to $\beta _{\phi }$ resulting in $\beta _{\phi } = 2.7 \pm 1.7$. After including a prior on the distortion parameters from constraints using Planck we find $\beta _{\phi } = 2.7^{+0.60}_{-0.67}$ suggesting $\beta _{\phi } > 0$ at 4.3$\sigma$ significance. This result may hint at a phase shift that is not purely sourced from the standard model expectation for $N_{\rm {eff}}$ or could be a upwards statistical fluctuation in the measured $\beta _{\phi }$; this result relaxes in models with additional freedom beyond Lambda-cold dark matter.

79 ASTRONOMY AND ASTROPHYSICS↗

Intern Poster

Large Language Models (LLMs) have skyrocketed in popularity after the release of ChatGPT in late 2022. Although LLMs are powerful tools, they can be subject to hallucinations, which is when an LLM (or any AI model) produces misleading/ nonsensical information. The objective is to determine if statistical methods can be used to detect hallucinations as an LLM generates its answer token by token (essentially word by word).

97 - MATHEMATICS AND COMPUTING↗

Insights into distorted lamellar phases with small-angle scattering and machine learning

Lamellar phases are essential in various soft matter systems, with topological defects significantly influencing their mechanical properties. In this report, we present a machine-learning approach for quantitatively analyzing the structure and dynamics of distorted lamellar phases using scattering techniques. By leveraging the mathematical framework of Kolmogorov–Arnold networks, we demonstrate that the conformations of these distorted phases – expressed as superpositions of complex waves – can be reconstructed from small-angle scattering intensities. Through the contour analysis of wave field phase singularities, we obtain the statistics of the spatial distribution of topological defects. Furthermore, we establish that the temporal evolution of these defects can be derived from the time-dependent traveling wave field, informed by the dispersion relation of spectral components. This method opens new avenues for investigating the dynamics of distorted lamellar phases using various dynamic scattering techniques such as neutron spin echo and X-ray photon correlation spectroscopy. These findings enhance our microscopic understanding of how defects influence the physical properties of lamellar materials, with implications for both equilibrium and non-equilibrium states in general lamellar systems.

36 MATERIALS SCIENCE↗

Multidimensional Distributional Neural Network Output Demonstrated in Super‐Resolution of Surface Wind Speed

Accurate quantification of uncertainty in neural network predictions remains a central challenge for scientific applications involving high-dimensional, correlated data. While existing methods capture either aleatoric or epistemic uncertainty, few offer closed-form, multidimensional distributions that preserve spatial correlation while remaining computationally tractable. In this work, we present a framework for training neural networks with a multidimensional Gaussian loss, generating a closed-form predictive distribution over outputs informed by non-identically distributed training data. Our approach captures aleatoric uncertainty by iteratively estimating the means and covariance matrices, and is demonstrated on a super-resolution example out-of-training-sample. We leverage a Fourier representation of the covariance matrix to stabilize network training and preserve spatial correlation. We introduce a novel regularization strategy—referred to as information sharing—that interpolates between image-specific and global covariance estimates, enabling convergence of the super-resolution downscaling network trained on image-specific distributional loss functions. This framework allows for efficient sampling, explicit correlation modeling, and extensions to more complex distribution families all without disrupting prediction performance. We demonstrate the method on a surface wind speed downscaling task and discuss its broader applicability to uncertainty-aware prediction in scientific models.

17 WIND ENERGY↗

Toward Accurate Spin–Orbit Splittings from Relativistic Multireference Electronic Structure Theory

Most nonrelativistic electron correlation methods can be adapted to account for relativistic effects, as long as the relativistic molecular spinor integrals are available, from either a four-, two-, or one-component mean-field calculation. Furthermore, relativistic multireference correlation methods remain a relatively unexplored area, with mixed evidence regarding the improvements brought by perturbative treatments. We report, for the first time, the implementation of state-averaged four-component relativistic multireference perturbation theories to second and third order based on the driven similarity renormalization group (DSRG). With our methods, named 4c-SA-DSRG-MRPT2 and 3, we find that the dynamical correlation included on top of 4c-CASSCF references can significantly improve the spin-orbit splittings in p-block elements and potential energy surfaces when compared to 4c-CASSCF and 4c-CASPT2 results. We further show that 4c-DSRG-MRPT2 and 3 are applicable to these systems over a wide range of the flow parameter, with systematic improvement from second to third order in terms of both improved error statistics and reduced sensitivity with respect to the flow parameter.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Overview of Electric Vehicle Load Modeling Strategies for Grid Integration Studies

The adoption of electric vehicles (EVs) has emerged as a solution to reduce greenhouse gas emissions in the transportation sector, which has motivated the implementation of public policies to promote their use in several countries. However, the high adoption of EVs poses challenges for the electricity sector, as it would imply an increase in energy demand and possible impacts on the power quality (PQ) of the power grid. Therefore, it is important to conduct EV integration studies in the power grid to determine the amount that can be incorporated without causing problems and identify the areas of the power sector that will require reinforcements. Accurate EV load patterns are required for this type of study that, through mathematical modeling, reflect both the dynamic behavior and the factors that influence the decision to recharge EVs. This article aims to present an overview of EVs, examine the different factors considered in the literature for modeling EV load patterns, and review modeling methods. EV load modeling methods are classified into deterministic, statistical, and machine learning. The article shows that each modeling method has its advantages, disadvantages, and data requirements, ranging from simple load modeling to more accurate models requiring large datasets.

Computer Science↗

Entropy Analysis of FPGA Interconnect and Switch Matrices for Physical Unclonable Functions

Random variations in microelectronic circuit structures represent the source of entropy for physical unclonable functions (PUFs). In this paper, we investigate delay variations that occur through the routing network and switch matrices of a field-programmable gate array (FPGA). The delay variations are isolated from other components of the programmable logic, e.g., look-up tables (LUTs), flip-flops (FFs), etc., using a feature of Xilinx FPGAs called dynamic partial reconfiguration (DPR). A set of partial designs is created to fix the placement of a time-to-digital converter (TDC) and supporting infrastructure to enable the path delays through the target interconnect and switch matrices to be extracted by subtracting out common-mode delay components. Delay variations are analyzed in the different levels of routing resources available within FPGAs, i.e., local routing and across-chip routing. Data are collected from a set of Xilinx Zynq 7010 devices, and a statistical analysis of within-die variations in delay through a set of the randomly-generated and hand-crafted interconnects is presented.

97 MATHEMATICS AND COMPUTING↗