Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Kinetic theory of particle-in-cell simulation plasma and the ensemble averaging technique

Abstract We derive the kinetic theory of fluctuations in physically and numerically stable particle-in-cell (PIC) simulations of electrostatic plasmas. The starting point is the single-time correlation at the start of the simulation between the statistical fluctuations of the weighted densities of macroparticle centers in the plasma particle phase-space. The fluctuations are associated with different initial conditions, typically due to the random initial conditions (in velocity space) of the macroparticles/simulation plasma, assigned according to their initial distribution of probability. The single-time correlations at all time steps and in each spatial grid cell are then determined from the Laplace–Fourier transforms of the discretized Klimontovich-like equation for the macroparticles and Maxwell’s equations for the fields, as computed by modern PIC codes. We recover the expressions for the electrostatic field and the plasma particle density fluctuation autocorrelation spectra as well as the kinetic equations describing the average evolution of PIC-simulated plasma particles, first derived by Langdon (1970b Proc. 4th Conf. Numerical Simulation of Plasmas ) using a test macroparticle approach perturbing a discretized Vlasovian plasma and then averaging the obtained physical quantity over the initial macroparticle velocity distribution. We generalize and extend these results to the modern algorithms in PIC codes using arbitrary macroparticle weights. Analytical estimates of statistical fluctuation amplitudes are derived as a function of the plasma simulation parameters, using the central limit theorem in the limit of a large number of macroparticles per cell. The theory is then used to analyze the ensemble averaging technique of PIC simulations where statistical averages are performed over ensembles of PIC simulations, modeling the same plasma physics problem but using different statistical realizations of the initial distribution functions of the macroparticles. This method is illustrated by linear Landau damping uncovering (from noise, which is usually considered numerical) the physical fluctuations driven by a single small amplitude electrostatic wave perturbing a PIC simulation plasma in equilibrium.

fluctuations correlations↗

Comparative investigations of multi-fidelity modeling on performance of electrostatically-actuated cracked micro-beams

Silicon is a commonly used material for the fabrication of beams for use in micro-electrical-mechanical systems (MEMS). Although silicon is a brittle material, it has been shown to accumulate fatigue damage at the micro-scale. Understanding the effect this has on the overall device performance is critical to the design of reliable devices. Analytical methods for modeling damage provide expedient results but are limited by broad modeling assumptions. Numerical models account for more detailed physical phenomena but can be computationally intensive. In this work, two different crack scenarios are modeled using both analytical techniques and 3D computational simulations. First, the effects of a single surface crack on the static deflection and natural frequency of an electrostatically actuated micro-beam are formulated and compared. Then, a new method for approximating damage associated with realistic distributed crack networks is formulated for use in an analytical model and numerical simulations. A method for utilizing experimentally derived crack statistics to inform the analytical and numerical distributed crack models is developed. Good agreement between the analytical and numerical models is obtained for both crack scenarios. Altogether, these models can be used to effectively simulate a variety of damage and fatigue behaviors in silicon-based MEMS devices.

42 ENGINEERING↗

Statistical Significance Testing for Mixed Priors: A Combined Bayesian and Frequentist Analysis

In many hypothesis testing applications, we have mixed priors, with well-motivated informative priors for some parameters but not for others. The Bayesian methodology uses the Bayes factor and is helpful for the informative priors, as it incorporates Occam’s razor via the multiplicity or trials factor in the look-elsewhere effect. However, if the prior is not known completely, the frequentist hypothesis test via the false-positive rate is a better approach, as it is less sensitive to the prior choice. We argue that when only partial prior information is available, it is best to combine the two methodologies by using the Bayes factor as a test statistic in the frequentist analysis. We show that the standard frequentist maximum likelihood-ratio test statistic corresponds to the Bayes factor with a non-informative Jeffrey’s prior. We also show that mixed priors increase the statistical power in frequentist analyses over the maximum likelihood test statistic. We develop an analytic formalism that does not require expensive simulations and generalize Wilks’ theorem beyond its usual regime of validity. In specific limits, the formalism reproduces existing expressions, such as the p-value of linear models and periodograms. We apply the formalism to an example of exoplanet transits, where multiplicity can be more than 10 7 . We show that our analytic expressions reproduce the $p$-values derived from numerical simulations. We offer an interpretation of our formalism based on the statistical mechanics. We introduce the counting of states in a continuous parameter space using the uncertainty volume as the quantum of the state. We show that both the $p$-value and Bayes factor can be expressed as an energy versus entropy competition.

97 MATHEMATICS AND COMPUTING↗

Statistical mechanical model for crack growth

Analytic relations that describe crack growth are vital for modeling experiments and building a theoretical understanding of fracture. Upon constructing an idealized model system for the crack and applying the principles of statistical thermodynamics, it is possible to formulate the rate of thermally activated crack growth as a function of load, but the result is analytically intractable. In this report an asymptotically correct theory is used to obtain analytic approximations of the crack growth rate from the fundamental theoretical formulation. These crack growth rate relations are compared to those that exist in the literature and are validated with respect to Monte Carlo calculations and experiments. The success of this approach is encouraging for future modeling endeavors that might consider more complicated fracture mechanisms, such as inhomogeneity or a reactive environment.

36 MATERIALS SCIENCE↗

Comparing galaxy formation in the L-GALAXIES semi-analytical model and the IllustrisTNG simulations

ABSTRACT We perform a comparison, object by object and statistically, between the Munich semi-analytical model, L-GALAXIES, and the IllustrisTNG hydrodynamical simulations. By running L-GALAXIES on the IllustrisTNG dark matter-only merger trees, we identify the same galaxies in the two models. This allows us to compare the stellar mass, star formation rate, and gas content of galaxies, as well as the baryonic content of subhaloes and haloes in the two models. We find that both the stellar mass functions and the stellar masses of individual galaxies agree to better than ${\sim} 0.2\,$dex. On the other hand, specific star formation rates and gas contents can differ more substantially. At z = 0, the transition between low-mass star-forming galaxies and high-mass quenched galaxies occurs at a stellar mass scale ${\sim} 0.5\,$dex lower in IllustrisTNG than that in L-GALAXIES. IllustrisTNG also produces substantially more quenched galaxies at higher redshifts. Both models predict a halo baryon fraction close to the cosmic value for clusters, but IllustrisTNG predicts lower baryon fractions in group environments. These differences are primarily due to differences in modelling feedback from stars and supermassive black holes. The gas content and star formation rates of galaxies in and around clusters and groups differ substantially, with IllustrisTNG satellites less star forming and less gas rich. We show that environmental processes such as ram-pressure stripping are stronger and operate to larger distances and for a broader host mass range in IllustrisTNG. We suggest that the treatment of galaxy evolution in the semi-analytic model needs to be improved by prescriptions that capture local environmental effects more accurately.

79 ASTRONOMY AND ASTROPHYSICS↗

Overview of the Tolerance Limit Calculations with Application to TSURFER

To establish confidence in the results of computerized physics models, a key regulatory requirement is to develop a scientifically defendable process. The methods employed for confidence, characterization, and consolidation, or C3, are statistically involved and are often accessible only to avid statisticians. This manuscript serves as a pedagogical presentation of the C3 process to all stakeholders—including researchers, industrial practitioners, and regulators—to impart an intuitive understanding of the key concepts and mathematical methods entailed by C3. The primary focus is on calculation of tolerance limits, which is the overall goal of the C3 process. Tolerance limits encode the confidence in the calculation results as communicated to the regulator. Understanding the C3 process is especially critical today, as the nuclear industry is considering more innovative ways to assess new technologies, including new reactor and fuel concepts, via an integrated approach that optimally combines modeling and simulation and minimal targeted validation experiments. This manuscript employs intuitive, analytical, numerical, and visual representations to explain how tolerance limits may be calculated for a wide range of configurations, and it also describes how their values may be interpreted. Various verification tests have been developed to test the calculated tolerance limits and to help delineate their values. The manuscript demonstrates the calculation of tolerance limits for TSURFER, a computer code developed by the Oak Ridge National Laboratory for criticality safety applications. The goal is to evaluate the tolerance limit for TSURFER-determined criticality biases to support the determination of upper, subcritical limits for regulatory purposes.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Evidence for galaxy assembly bias in BOSS CMASS redshift-space galaxy correlation function

ABSTRACT Building accurate and flexible galaxy–halo connection models is crucial in modelling galaxy clustering on non-linear scales. Recent studies have found that halo concentration by itself cannot capture the full galaxy assembly bias effect and that the local environment of the halo can be an excellent indicator of galaxy assembly bias. In this paper, we propose an extended halo occupation distribution (HOD) model that includes both a concentration-based assembly bias term and an environment-based assembly bias term. We use this model to achieve a good fit (χ2/degrees of freedom = 1.35) on the 2D redshift-space two-point correlation function (2PCF) of the Baryon Oscillation Spectroscopic Survey (BOSS) CMASS galaxy sample. We find that the inclusion of both assembly bias terms is strongly favoured by the data and the standard five-parameter HOD model is strongly rejected. More interestingly, the redshift-space 2PCF drives the assembly bias parameters in a way that preferentially assigns galaxies to lower mass haloes. This results in galaxy–galaxy lensing predictions that are within 1σ agreement with the observation, alleviating the perceived tension between galaxy clustering and lensing. We also showcase a consistent 3σ–5σ preference for a positive environment-based assembly bias that persists over variations in the fit. We speculate that the environmental dependence might be driven by underlying processes such as mergers and feedback, but might also be indicative of a larger halo boundaries such as the splashback radius. Regardless, this work highlights the importance of building flexible galaxy–halo connection models and demonstrates the extra constraining power of the redshift-space 2PCF.

79 ASTRONOMY AND ASTROPHYSICS↗

Hydrodynamic fluctuations near a Hopf bifurcation: Stochastic onset of vortex shedding behind a circular cylinder

Here, we investigate hydrodynamic fluctuations in the flow past a circular cylinder near the critical Reynolds number Re c for the onset of vortex shedding. Starting from the fluctuating Navier-Stokes equations, we perform a perturbation expansion around Re c to derive analytical expressions for the statistics of the fluctuating lift force. Molecular-level simulations using the direct simulation Monte Carlo method support the theoretical predictions of the lift power spectrum and amplitude distribution. Notably, we have been able to collect sufficient statistics at distances Re ⁡/ Re c – 1 = O ⁡(10 –3 ) from the instability that confirm the appearance of non-Gaussian fluctuations, and we observe that they are associated with intermittent vortex shedding. These results emphasize how unavoidable thermal-noise-induced fluctuations become dramatically amplified in the vicinity of oscillatory flow instabilities and that their onset is fundamentally stochastic.

42 ENGINEERING↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

Artificial Intelligence/Machine Learning Technologies for Advanced Reactors (Workshop Summary Report)

A workshop on artificial intelligence and machine learning (AI/ML) for advanced reactors (AR) was held October 5-6, 2021. The workshop was to be attended in-person at ANL but COVID restrictions forced the workshop to go virtual. The objectives of the workshop were to identify the most promising AI/ML opportunities for improving advanced reactor design, optimizing plant performance, and enhancing economic competitiveness and to develop an understanding of the scientific, engineering and licensing challenges facing their application. The workshop planning committee included GAIN, EPRI and NEI and members of three national laboratories (ANL, INL, and ORNL). The workshop was attended by more than 200 individuals representing academic and scientific institutions and the nuclear power industry. The definition put forth for an AI/ML system was one that perceives its environment and takes actions that maximize its chance of achieving its goals. In this report AI/ML refers to next generation algorithms that include deep learning, statistical analysis and data analytics and associated scientific computing and their potential application to the design, licensing, operation and maintenance of ARs. These methods typically incorporate models built from process data and may also include data generated by simulations that represent the behavior of a system. The workshop was organized in response to the growing interest in application of AI/ML for improving the economic competitiveness of nuclear energy. Increasingly more resources are being allocated to investigating the benefits of AI/ML methods. The DOE created the Artificial Intelligence & Technology Office to promote their development. And within the Office of Nuclear Energy, resources have been allocated to explore and understand the potential benefits of AI/ML. Additionally, the national laboratories are strategically positioned with DOE computing facilities such as Summit, Perlmutter, Aurora and Frontier that support large-scale simulations, hybrid HPC models with AI surrogates, and the exploration of new types of generative models emerging from multi-model data streams and sources. The workshop was organized with members of the AR community to understand the effort and to identify the level of interest and progress in this emerging technology. The workshop discussions focused on identifying opportunities for AI/ML across diverse areas of the nuclear industry and identifying current scientific and engineering challenges for advanced reactors that might be addressed through transformational uses of AI/ML. Discussion panels focused on four high-interest technical domains for advanced reactors: design, maintenance and operations, energy storage, and materials. The results of those discussions are summarized in this report. This includes opportunities that were identified for exploiting AI techniques and methods to improve the efficacy and efficiency of reactor analysis and to improve the operation and optimization of advanced reactors. Advanced reactor developers expressed an interest in learning more about AI/ML methods and their application. This included understanding whether ML methods can provide an advantage over existing nonlinear data regression methods for collapsing high-fidelity simulation results into faster running models. A consensus emerged that AR advances planned for the next decade will benefit from the use of AI/ML tools. The need exists to understand and model complex systems across length scales and modalities. AI/ML is a tool for discovery that can yield a set of engineering principles for use by nuclear engineers, licensing bodies, and operators to solve problems in plant design, safety analyses, autonomous operation, and predictive maintenance. While AI/ML represents a new set of tools, an awareness by the nuclear community of the full potential is still in the early stages so there is a need to increase awareness. It appears that the wide-spread adoption of AI/ML tools for ARs would be facilitated by future educational workshops that describe foundational methods and capabilities and describe successful applications.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Chaconne: A Statistical Approach to Nonlocal Compression for Supervised Learning, Semi-Supervised Learning, and Anomaly Detection

This project developed a novel statistical understanding of compression analytics (CA), which has challenged and clarified some core assumptions about CA, and enabled the development of novel techniques that address vital challenges of national security. Specifically, this project has yielded the development of novel capabilities including 1. Principled metrics for model selection in CA, 2. Techniques for deriving/applying optimal classification rules and decision theory to supervised CA, including how to properly handle class imbalance and differing costs of misclassification, 3. Two techniques for handling nonlocal information in CA, 4. A novel technique for unsupervised CA that is agnostic with regard to the underlying compression algorithm, 5. A framework for semisupervised CA when a small number of labels are known in an otherwise large unlabeled dataset. 6. The academic alliance component of this project has focused on the development of a novel exemplar-based Bayesian technique for estimating variable length Markov models (closely related to PPM [prediction by partial matching] compression techniques). We have developed examples illustrating the application of our work to text, video, genetic sequences, and unstructured cybersecurity log files.

99 GENERAL AND MISCELLANEOUS↗

A New Window into the Baryon Cycle at Cosmic Noon with Line Intensity Mapping: Forecasts for auto- and cross-correlations in [CII]-158$μ$m, HI 21 cm, CO$_{J+1\rightarrow J}$, and H$α$ galaxies

Across the peak of cosmic star formation at $z\sim1-2$, inflow, processing, and feedback drive rapid changes in the spatial distribution and chemical composition of baryons in galaxies and surrounding reservoirs; this baryon cycle can be tomographically mapped by line intensity mapping (LIM) of atomic hydrogen, ionized carbon, and carbon monoxide. We present a simulation-based forecasting framework for detecting auto- and cross-power spectra between spectroscopic surveys of four such tracers at $z\sim0.5-1.7$ mapping the same deep field - TIM, EoRSpec/FYST, MeerKAT, & Euclid. We forward-model 3-D distributions for these tracers from magnetohydrodynamic simulations, directly capturing the two-halo, one-halo, and shot statistics without relying on analytical decompositions. We further detail a signal-to-noise formalism, tailored to LIM surveys with highly anisotropic geometries and Fourier-space coverage. We demonstrate that galaxy cross-correlations will be the dominant discovery channel for current-generation surveys. These instruments will detect the auto-spectra for CO and HI 21 cm and the CO $\times$ 21 cm cross-spectrum at modest S/N $\sim 1-10$, while placing upper limits on the [CII]-158$μ$m signals. [CII], CO, and HI LIM will be $\sim3-30\times$ ($0.5-1.5$ dex) more sensitive to cross-correlation with the Euclid survey, however, than their respective auto-correlations, constraining all three models of line emission at high significance (S/N $\sim 10-40$) within this decade. Finally, we formulate a staged instrumental trajectory with planned or reasonable improvements, including the as-proposed SKA-Mid. We forecast advancing the per-$k$-mode sensitivities of each auto-, galaxy-line, and line-line spectrum by several orders of magnitude, enabling new percent- and sub-percent level constraints on cosmology and the redshift evolution of star formation and the baryon cycle.

Agrawal, Shubh [Pennsylvania U., Dept. Math.]↗

Streaming Statistics

In the context of a larger effort for in situ data analytics, there is a need to calculate basic statistics metrics (e.g., count, mean, median) online as new data points become available. Originally, the code for such online, or in other words streaming, statistics was part of the TALASS (Topological Analysis of Large- Scale Simulations) library. We isolated the relevant code and created a standalone library from it called Streaming Statistics. We also added an ability to serialize and deserialize the statistics objects so that the library can be used in distributed, task-based processing. To use the Streaming Statistics library, the user chooses a statistic, constructs an object for it, and then "adds" values to it, which means the statistic is augmented.

Shudler, Sergei↗

Hypothesis-Agnostic Network-Based Analysis of Real-World Data Suggests Ondansetron is Associated with Lower COVID-19 Any Cause Mortality

Background: The COVID-19 pandemic generated a massive amount of clinical data, which potentially hold yet undiscovered answers related to COVID-19 morbidity, mortality, long-term effects, and therapeutic solutions.Objectives: The objectives of this study were (1) to identify novel predictors of COVID-19 any cause mortality by employing artificial intelligence analytics on real-world data through a hypothesis-agnostic approach and (2) to determine if these effects are maintained after adjusting for potential confounders and to what degree they are moderated by other variables.Methods: A Bayesian statistics-based artificial intelligence data analytics tool (bAIcis®) within the Interrogative Biology® platform was used for Bayesian network learning and hypothesis generation to analyze 16,277 PCR+ patients from a database of 279,281 inpatients and outpatients tested for SARS-CoV-2 infection by antigen, antibody, or PCR methods during the first pandemic year in Central Florida. This approach generated Bayesian networks that enabled unbiased identification of significant predictors of any cause mortality for specific COVID-19 patient populations. These findings were further analyzed by logistic regression, regression by least absolute shrinkage and selection operator, and bootstrapping.Results: We found that in the COVID-19 PCR+ patient cohort, early use of the antiemetic agent ondansetron was associated with decreased any cause mortality 30 days post-PCR+ testing in mechanically ventilated patients.Conclusions: The results demonstrate how a real-world COVID-19-focused data analysis using artificial intelligence can generate unexpected yet valid insights that could possibly support clinical decision making and minimize the future loss of lives and resources.

60 APPLIED LIFE SCIENCES↗

Spectral form factor of a quantum spin glass

It is widely expected that systems which fully thermalize are chaotic in the sense of exhibiting random-matrix statistics of their energy level spacings, whereas integrable systems exhibit Poissonian statistics. In this paper, we investigate a third class: spin glasses. These systems are partially chaotic but do not achieve full thermalization due to large free energy barriers. We examine the level spacing statistics of a canonical infinite-range quantum spin glass, the quantum p-spherical model, using an analytic path integral approach. We find statistics consistent with a direct sum of independent random matrices, and show that the number of such matrices is equal to the number of distinct metastable configurations — the exponential of the spin glass “complexity” as obtained from the quantum Thouless-Anderson-Palmer equations. We also consider the statistical properties of the complexity itself and identify a set of contributions to the path integral which suggest a Poissonian distribution for the number of metastable configurations. Our results show that level spacing statistics can probe the ergodicity-breaking in quantum spin glasses and provide a way to generalize the notion of spin glass complexity beyond models with a semi-classical limit.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Benchmark Solutions for Radiation Transport in Stochastic Media with Inhomogeneous Material Statistics

Accurately solving implicit Monte Carlo (IMC) thermal photon transport problems with mixed material cells is important in realistic applications. The production IMC package at LLNL treats mixed material cells arising from ALE remap and hydrodynamics using the same approximate model. The new Imp IMC thermal photon transport package currently under development has both a material interface reconstruction (MIR) algorithm and a Levermore-Pomraning (LP) stochastic medium algorithm for treating mixed material cells. Existing stochastic medium algorithms for treating mixed material cells in IMC lack a complete theoretical basis. The IMC LP algorithm implementation has been demonstrated to reproduce published deterministic LP solutions for the particular case of spatially homogeneous material statistics. Realistic simulations will include spatially inhomogeneous material statistics (material mean chord lengths). In a previous investigation, the LP-model for transport in binary stochastic media in rod geometry was generalized to accommodate spatially varying material chord lengths, i.e., the mixing statistics were allowed to be nonhomogeneous. Analytical solutions were obtained and used to produce a verifi cation suite for the Imp IMC Levermore-Pomraning implementation for different spatial variations of the chord lengths. However, the accuracy of the LP model when the mixing statistics are nonhomogeneous has not been assessed and leaves open the question of whether local accuracy is improved or further degraded when chord lengths are not uniform. This shortcoming is rectifi ed here by developing benchmark analytic solutions for transport in binary Markovian stochastic mixtures in rod geometry with nonhomogeneous mixing statistics, using spatially varying chord lengths considered in the previous investigation based on the LP model. Methods for sampling a nonhomogeneous Poisson process (NHPP) are first described and used to construct individual realizations of the binary mixtures in rod geometry. Analytic solutions are then obtained for the forward and backward directed fluxes on a given realization, now viewed as a deterministic medium with alternating layers of the two materials with known interface locations. Finally, material averaged scalar fluxes are obtained using these sampling schemes with spatially linear and quadratic chord lengths and used to assess the accuracy of the previously obtained LP-model results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Research Needs for Trusted Analytics in National Security Settings

As artificial intelligence, machine learning, and statistical modeling methods become commonplace in national security applications, the drive to create trusted analytics becomes increasingly important. The goal of this report is to identify areas of research that can provide the foundational understanding and technical prerequisites for the development and deployment of trusted analytics in national security settings. Our review of the literature covered several disjoint research communities, including computer science, statistics, human factors, and several branches of psychology and cognitive science, which tend not to interact with one another or cite each other's literatures. As a result, there exists no agreed-upon theoretical framework for understanding how various factors influence trust and no well-established empirical paradigm for studying these effects. This report therefore takes three steps. First, we define several key terms in an effort to provide a unifying language for trusted analytics and to manage the scope of the problem. Second, we outline an empirical perspective that identifies key independent, moderating, and dependent variables in assessing trusted analytics. Though not a substitute for a theoretical framework, the empirical perspective does support research and development of trusted analytics in the national security domain. Finally, we discuss several research gaps relevant to developing trusted analytics for the national security mission space.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Toward Accurate Modeling of Galaxy Clustering on Small Scales: Constraining the Galaxy-halo Connection with Optimal Statistics

Applying halo models to analyze the small-scale clustering of galaxies is a proven method for characterizing the connection between galaxies and their host halos. Such works are often plagued by systematic errors or limited to clustering statistics that can be predicted analytically. In this work, we employ a numerical mock-based modeling procedure to examine the clustering of Sloan Digital Sky Survey DR7 galaxies. We apply a standard halo occupation distribution (HOD) model to dark matter only simulations with a ΛCDM cosmology. To constrain the theoreStical models, we utilize a combination of galaxy number density and selected scales of the projected correlation function, redshift-space correlation function, group multiplicity function, average group velocity dispersion, mark correlation function, and counts-in-cells statistics. We design an algorithm to choose an optimal combination of measurements that yields tight and accurate constraints on our model parameters. Compared to previous work using fewer clustering statistics, we find a significant improvement in the constraints on all parameters of our halo model for two different luminosity-threshold galaxy samples. Most interestingly, we obtain unprecedented high-precision constraints on the scatter in the relationship between galaxy luminosity and halo mass. However, our best-fit model results in significant tension (>4σ) for both samples, indicating the need to add second-order features to the standard HOD model. To guarantee the robustness of these results, we perform an extensive analysis of the systematic and statistical errors in our modeling procedure, including a first of its kind study of the sensitivity of our constraints to changes in the halo mass function due to baryonic physics.

79 ASTRONOMY AND ASTROPHYSICS↗