Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Optimization of the deep neural network parameters for generating homogenized fuel assembly data for nodal codes

Homogenized fuel assembly (FA) data is a typical input data for nodal codes. Generating that data, however, could be time-consuming. One of promising ways to mitigate the computational burden of generating macroscopic cross-sections is to use trained artificial neural network (ANN) models for predicting nuclear data. However, there is a challenge to make the model support variable FA geometry. In this work, two most common types of FA were combined in one ANN model. Since there could be multiple ways of converting 2-dimensional FA data into 1-dimensional input vector for ANN, three different approaches of data flattening were evaluated. The input parameters included each fuel pin enrichment, fuel temperature, moderator temperature and boron concentration. The output parameters were 2-group macroscopic cross-sections (XS) and pin power distribution (HFF). A fully connected deep neural network (DNN) model was trained and tested using pre-generated data obtained with lattice physics code STREAM. The results of this study showed no statistically significant difference in the accuracy of XS and HFF generation for all 3 tested input vector orders. This means that fully connected DNN for XS generation demonstrated input sequence invariance. Results of comparing predicted XS data with reference solutions were found sufficiently close considering the reduction of computation time offered by ANN. Mean relative difference (MRD) for all output XS parameters was found below 0.7%, while HFF MRD was found higher compared to XS values, in some cases slightly exceeding 1%, mostly near guide tube locations. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

DESI mock challenge: Halo and galaxy catalogues with the bias assignment method

We present a novel approach to the construction of mock galaxy catalogues for large-scale structure analysis based on the distribution of dark matter halos obtained with effective bias models at the field level. We aim to produce mock galaxy catalogues capable of generating accurate covariance matrices for a number of cosmological probes that are expected to be measured in current and forthcoming galaxy redshift surveys (e.g. two- and three-point statistics). The construction of the catalogues shown in this paper is part of a mock-comparison project within the Dark Energy Spectroscopic Instrument (DESI) collaboration. We use the bias assignment method ( BAM ) to model the statistics of halo distribution through a learning algorithm using a few detailed N-body simulations, and approximated gravity solvers based on Lagrangian perturbation theory. We introduce cosmic-web-dependent corrections to modelling redshift-space distortions at the N-body level – both in the halo and galaxy distributions –, as well as a multi-scale approach for accurate assignment of halo properties. Using specific models of halo occupation distributions to populate halos, we generate galaxy mocks with the expected number density and central-satellite fraction of emission-line galaxies, which are a key target of the DESI experiment. BAM generates mock catalogues with per cent accuracy in a number of summary statistics, such as the abundance, the two- and three-point statistics of halo distributions, both in real and redshift space. In particular, the mock galaxy catalogues display ~3%-10% accuracy in the multipoles of the power spectrum up to scales of k ~ 0.4 h -1 Mpc. We show that covariance matrices of two- and three-point statistics obtained with BAM display a similar structure to the reference simulation. BAM offers an efficient way to produce mock halo catalogues with accurate two- and three-point statistics and is able to generate a variety of multi-tracer catalogues with precise covariance matrices of several cosmological probes. We discuss future developments of the algorithm towards mock production in DESI and other galaxy-redshift surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark Energy Survey: Modeling strategy for multiprobe cluster cosmology and validation for the Full Six-year Dataset

We introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3x2pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (1) identification of modeling and parameterization choices and impact studies using simulated analyses with each possible model misspecification (2) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3x2pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy--galaxy lensing and cosmic shear two-point statistics) and the CL+3x2pt+BAO+SN (combination of CL+3x2pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dark energy survey: Modeling strategy for multiprobe cluster cosmology and validation for the full six-year dataset

Here, we introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3×2 pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (i) identification of modeling and parametrization choices and impact studies using simulated analyses with each possible model misspecification and (ii) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3×2 pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy–galaxy lensing and cosmic shear two-point statistics) and the CL+3×2 pt+BAO+SN (combination of CL+3×2 pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

79 ASTRONOMY AND ASTROPHYSICS↗

Nonlinear encoding in diffractive information processing using linear optical materials

Nonlinear encoding of optical information can be achieved using various forms of data representation. Here, we analyze the performances of different nonlinear information encoding strategies that can be employed in diffractive optical processors based on linear materials and shed light on their utility and performance gaps compared to the state-of-the-art digital deep neural networks. For a comprehensive evaluation, we used different datasets to compare the statistical inference performance of simpler-to-implement nonlinear encoding strategies that involve, e.g., phase encoding, against data repetition-based nonlinear encoding strategies. We show that data repetition within a diffractive volume (e.g., through an optical cavity or cascaded introduction of the input data) causes the loss of the universal linear transformation capability of a diffractive optical processor. Therefore, data repetition-based diffractive blocks cannot provide optical analogs to fully connected or convolutional layers commonly employed in digital neural networks. However, they can still be effectively trained for specific inference tasks and achieve enhanced accuracy, benefiting from the nonlinear encoding of the input information. Our results also reveal that phase encoding of input information without data repetition provides a simpler nonlinear encoding strategy with comparable statistical inference accuracy to data repetition-based diffractive processors. Our analyses and conclusions would be of broad interest to explore the push-pull relationship between linear material-based diffractive optical systems and nonlinear encoding strategies in visual information processors.

42 ENGINEERING↗

Accurate core excitation and ionization energies from a state-specific coupled-cluster singles and doubles approach

Here, we investigate the use of orbital-optimized references in conjunction with single-reference coupled-cluster theory with single and double substitutions (CCSD) for the study of core excitations and ionizations of 18 small organic molecules, without the use of response theory or equation-of-motion (EOM) formalisms. Three schemes are employed to successfully address the convergence difficulties associated with the coupled-cluster equations, and the spin contamination resulting from the use of a spin symmetry-broken reference, in the case of excitations. In order to gauge the inherent potential of the methods studied, an effort is made to provide reasonable basis set limit estimates for the transition energies. Overall, we find that the two best-performing schemes studied here for ΔCCSD are capable of predicting excitation and ionization energies with errors comparable to experimental accuracies. The proposed ΔCCSD schemes reduces statistical errors against experimental excitation energies by more than a factor of two when compared to the frozen-core core–valence separated (FC-CVS) EOM-CCSD approach – a successful variant of EOM-CCSD tailored towards core excitations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Wavelet flow for extragalactic foreground simulations

Extragalactic foregrounds in cosmic microwave background (CMB) observations are both a source of cosmological and astrophysical information and a nuisance to the CMB. Effective field-level modeling that captures their non-Gaussian statistical distributions is increasingly important for optimal information extraction, particularly given the low-noise observations from current and upcoming experiments. Here, we explore the use of Wavelet Flow (WF) models to tackle the novel task of modeling the field-level probability distributions of multi-component CMB secondaries and foregrounds. Specifically, we jointly train correlated CMB lensing convergence (κ) and cosmic infrared background (CIB) maps with a WF model and obtain a network that statistically recovers the input to high accuracy — the trained network generates samples of κ and CIB fields whose average power spectra are within a few percent of the inputs across all scales, and whose Minkowski functionals are similarly accurate compared to the inputs. Leveraging the multiscale architecture of these models, we fine-tune both the model parameters and the priors at each scale independently, optimizing performance across different resolutions. These results demonstrate that WF models can accurately simulate correlated components of CMB secondaries, supporting improved analysis of cosmological data. Our code and trained models can be found on this GitHub repo.

cosmological simulations↗

Exploiting non-linear scales in galaxy–galaxy lensing and galaxy clustering: A forecast for the dark energy survey

ABSTRACT The combination of galaxy–galaxy lensing (GGL) and galaxy clustering is a powerful probe of low-redshift matter clustering, especially if it is extended to the non-linear regime. To this end, we use an N-body and halo occupation distribution (HOD) emulator method to model the redMaGiC sample of colour-selected passive galaxies in the Dark Energy Survey (DES), adding parameters that describe central galaxy incompleteness, galaxy assembly bias, and a scale-independent multiplicative lensing bias Alens. We use this emulator to forecast cosmological constraints attainable from the GGL surface density profile ΔΣ(rp) and the projected galaxy correlation function wp, gg(rp) in the final (Year 6) DES data set over scales $r_p=0.3\!-\!30.0\, h^{-1} \, \mathrm{Mpc}$. For a $3{{\ \rm per\ cent}}$ prior on Alens we forecast precisions of $1.9{{\ \rm per\ cent}}$, $2.0{{\ \rm per\ cent}}$, and $1.9{{\ \rm per\ cent}}$ on Ωm, σ8, and $S_8 \equiv \sigma _8\Omega _m^{0.5}$, marginalized over all halo occupation distribution (HOD) parameters as well as Alens. Adding scales $r_p=0.3\!-\!3.0\, h^{-1} \, \mathrm{Mpc}$ improves the S8 precision by a factor of ∼1.6 relative to a large scale ($3.0\!-\!30.0\, h^{-1} \, \mathrm{Mpc}$) analysis, equivalent to increasing the survey area by a factor of ∼2.6. Sharpening the Alens prior to $1{{\ \rm per\ cent}}$ further improves the S8 precision to $1.1{{\ \rm per\ cent}}$, and it amplifies the gain from including non-linear scales. Our emulator achieves per cent-level accuracy similar to the projected DES statistical uncertainties, demonstrating the feasibility of a fully non-linear analysis. Obtaining precise parameter constraints from multiple galaxy types and from measurements that span linear and non-linear clustering offers many opportunities for internal cross-checks, which can diagnose systematics and demonstrate the robustness of cosmological results.

79 ASTRONOMY AND ASTROPHYSICS↗

Beyond the Hype: An Evaluation of Commercially Available Machine-Learning-Based Malware Detectors

There is a lack of scientific testing of commercially available malware detectors, especially those that boast accurate classification of never-before-seen (i.e., zero-day) files using machine learning (ML). Consequently, efficacy of malware detectors is opaque, inhibiting end users from making informed decisions and researchers from targeting gaps in current detectors. In this paper, we present a scientific evaluation of four prominent commercial malware detection tools to assist an organization with two primary questions: To what extent do ML-based tools accurately classify previously and never-before-seen files? Is purchasing a network-level malware detector worth the cost? To investigate, we tested each tool against 3,536 total files (2,554 or 72% malicious, 982 or 28% benign) of a variety of file types, including hundreds of malicious zero-days, polyglots, and APT-style files, delivered on multiple protocols. We present statistical results on detection time and accuracy, consider complementary analysis (using multiple tools together), and provide two novel applications of the recent cost-benefit evaluation procedure of Iannacone & Bridges. Although the ML-based tools are more effective at detecting zero-day files and executables, the signature-based tool might still be an overall better option. Both network-based tools provide substantial (simulated) savings when paired with either host tool, yet both show poor detection rates on protocols other than HTTP or SMTP. Our results show that all four tools have near-perfect precision but alarmingly low recall, especially on file types other than executables and office files—37% of malware, including all polyglot files, were undetected. Priorities for researchers and takeaways for end users are given. Code for future use of the cost model is provided.

97 MATHEMATICS AND COMPUTING↗

FutureTense

Protective vaccines and reliable diagnostics are essential tools for controlling viral diseases. However, the efficacy of these tools can be diminished by mutations in viral genomes. The delay between the emergence of new viral strains and the redesign of vaccines and diagnostics allows for continued viral transmission. Is it possible to address this challenge by computationally predicting viral genome sequence evolution? Can we “future-proof” vaccines and diagnostics by targeting both current and anticipated future sequence variants? While predicting viral evolution is still an unsolved, “grand challenge” problem in biology, the large, and rapidly growing, number of SARS-CoV-2 genome sequences provide an opportunity to quantify the ability of machine learning to predict viral genome sequence evolution. Towards this end, we have developed a simple computational model for predicting viral evolution at the level of individual nucleotides. The key metric for quantifying the per-base, prediction accuracy for viral evolution is the Mann-Whitney U statistic (or, equivalently, the area under the receiver operator curve). Since the Mann-Whitney U statistic is not a differentiable function, existing deep leaning packages (like Pytorch and Keras/TensorFlow) are not useful, as they require that the accuracy metric/objective function be analytically differentiable with respect to the model parameters. To overcome this challenge, we have implemented custom software, “FutureTense”, that can train a machine learning model by maximizing the non-differentiable Mann-Whitney U statistic. This software trains a machine learning model by exploring along the direction of the discrete gradient of the Mann-Whitney U statistic in the model parameter space. Parallel computing and genome sequence-specific optimizations are used to accelerate model training. The resulting machine learning model learns the observed high C->U mutation rates in the SARS-CoV-2 genome (which are potentially induced by host defenses) and provides prediction accuracies that are significantly better than one would expect from random chance. While predicting viral evolution is still quite far from a solved problem, the surprising performance of this simple model gives hope that the accuracy of predicting viral genome evolution can be further increased by more sophisticated approaches.

Gans, Jason↗

Extrapolation of the Rainflow-Counted Load Ranges for Fatigue Assessment of the Wind Turbine's Blades

Wind turbine design standards recommend the use of statistical modeling coupled with extrapolation of the short-term load data to long-term periods for fatigue reliability assessment. However, statistical error and computational expense can limit the accuracy of such approaches. In the case of wind turbine blades, the errors are more significant because of the high material fatigue exponent that makes the damage estimations more sensitive to variations. In addition, due to different excitation sources, the flapwise load range histogram is not unimodal, and thus its statistical modeling is complex. In the present work, we provide three methods for statistical modeling of the flapwise bending moment ranges including a novel approach based on frequency-based separation of the modes. The first two methods are simplified approaches for modeling the most crucial load ranges using unimodal distributions and the third method involves multimodal distribution fitting. The research is based on 3600 10-minute aeroelastic simulations of DTU 10MW case study wind turbine from which a benchmark damage equivalent load (DEL) is calculated. The DEL calculated by each of the three proposed methods is compared to this reference. The results show that the conventional approach based on using 6 seeds as well as using mixture models fitted on the limited data lead to under-conservative results with errors up to 23%. On the other hand, the simplified unimodal approaches provided in this work can provide conservative estimations of the fatigue damage with mean values 5% and 12% higher than the benchmark. However, the variability of the DEL estimates is higher when using unimodal extrapolation of the load ranges, and the data can be conservative by 17.5%. The proposed unimodal fits suggested for modeling and extrapolation of the blade's load ranges provide less errors relatively and most importantly conservative DEL estimations while maintaining computational efficiency.

blade fatigue↗

Constitutive model development of aluminum alloy 1100 for elevated temperature forming process

Commercially pure aluminum alloy, AA1100, presents good electrical and thermal conductivity, high formability, and low cost. Those favorable characteristics have the potential to enable bipolar plates with improved economics and enhanced performance compared to current stainless steel bipolar plates for proton exchange membrane fuel cells. An accurate constitutive model is essential to develop and optimize processing parameters and effectively control the forming process. Here, the objective of this work is to develop a constitutive model of AA1100 that is able to simulate stress-strain relation, formed geometry, and predict the onset of fracture strain to avoid forming failure. Initially, a set of tensile tests at temperature between 300 and 500°C and strain rate between 0.005 and 1.0/s were conducted to examine the deformation behavior. Then, a set of damage-based unified visco-plastic constitutive equations is proposed and calibrated based on the results of stress-strain data. A genetic algorithm optimization method is applied to search for best fitting material constants in constitutive equations. The proposed model shows good predictability of both the stress-strain relation and fracture strain at low strain rate and high temperature conditions. The accuracy of proposed model is also evaluated statistically. A comparison of the proposed model with three popular models (Arrhenius-type mode, Johnson-Cook model and Zerilli-Armstrong model) was made. The proposed model shows the best experimental agreement with correlation coefficient of 0.96 in contrast to 0.25, 0.38 and 0.75 for the popular models, respectively. The proposed model can help to optimize the elevated temperature forming process and guide die design to enable optimal geometric features in the formed components.

08 HYDROGEN↗

On the measurement of shape: With applications to lunar regolith

With the renewed commitment from NASA and other commercial entities for a presence on the Moon, the importance of understanding the characteristics of lunar regolith and how to utilize it have become the target of increasing scrutiny. Much of what is known about lunar regolith was collected during and immediately after the Apollo program, however, analytical techniques and instrumentation have advanced in leaps and bounds in the subsequent decades. Specifically, dynamic image analysis systems have advanced to the point that millions of particles can have morphological characteristics automatically determined in relatively short time frames. Particle morphological data was collected on several lunar samples and a pair of widely used regolith simulants to ascertain the accuracy of these simulants and to explore statistical analysis methods of these large datasets. It is found that these morphology datasets can vary widely depending on the particle size of the particles, and simple averaging of the data skews the results heavily towards the numerically abundant size ranges, the fines. Different reporting methods are suggested to ameliorate these problems. When applied to the lunar regolith, the particles are noted to be less morphologically complex than initially suspected. Compared to the lunar material, the simulants are found to contain some more morphological variability and angular grains. Such difference is likely due to the wildly different comminution processes that these different powder systems are subjected to.

2D shape↗

Accelerating phase field simulations through a hybrid adaptive Fourier neural operator with U-net backbone

Prolonged contact between a corrosive liquid and metal alloys can cause progressive dealloying. For one such process as liquid-metal dealloying (LMD), phase field models have been developed to understand the mechanisms leading to complex morphologies. However, the LMD governing equations in these models often involve coupled non-linear partial differential equations (PDE), which are challenging to solve numerically. In particular, numerical stiffness in the PDEs requires an extremely refined time step size (on the order of 10 -12 s or smaller). This computational bottleneck is especially problematic when running LMD simulation until a late time horizon is required. This motivates the development of surrogate models capable of leaping forward in time, by skipping several consecutive time steps at-once. In this paper, we propose a U-shaped adaptive Fourier neural operator (U-AFNO), a machine learning (ML) based model inspired by recent advances in neural operator learning. U-AFNO employs U-Nets for extracting and reconstructing local features within the physical fields, and passes the latent space through a vision transformer (ViT) implemented in the Fourier space (AFNO). We use U-AFNOs to learn the dynamics of mapping the field at a current time step into a later time step. We also identify global quantities of interest (QoI) describing the corrosion process (e.g., the deformation of the liquid-metal interface, lost metal, etc.) and show that our proposed U-AFNO model is able to accurately predict the field dynamics, in spite of the chaotic nature of LMD. Most notably, our model reproduces the key microstructure statistics and QoIs with a level of accuracy on par with the high-fidelity numerical solver, while achieving a significant 11, 200 × speed-up on a high-resolution grid when comparing the computational expense per time step. Finally, we also investigate the opportunity of using hybrid simulations, in which we alternate forward leaps in time using the U-AFNO with high-fidelity time stepping. We demonstrate that while advantageous for some surrogate model design choices, our proposed U-AFNO model in fully auto-regressive settings consistently outperforms hybrid schemes.

36 MATERIALS SCIENCE↗

Simulation budgeting for hybrid effective field theories

In this work, we forecast the number of, and requirements on, N-body simulations needed to train hybrid effective field theory (HEFT) emulators for a range of use cases, using a hybrid of HMcode and perturbation theory as a surrogate model. Our accuracy goals, determined with careful consideration of statistical and systematic uncertainties, are 1% accurate in the high-likelihood range of cosmological parameters, and 2% accurate over a broader parameter space volume for k < 1 h Mpc -1 and z < 3. Focusing in part on the 8-parameter w 0 w a CDM+ m ν cosmological model, we find that < 225 simulations are required to meet our error goals over our wide parameter space, including models with rapidly evolving dark energy, given our simulation and emulator recommendations. For a more restricted parameter space volume, as few as 80 simulations are sufficient. We additionally present simulation forecasts for example use cases, and make the code used in our analyses publicly available. These results offer practical guidance for efficient emulator design and simulation budgeting in future cosmological analyses.

cosmological parameters from LSS↗

A better way to define dark matter haloes

ABSTRACT Dark matter haloes have long been recognized as one of the fundamental building blocks of large-scale structure formation models. Despite their importance – or perhaps because of it! – halo definitions continue to evolve towards more physically motivated criteria. Here, we propose a new definition that is physically motivated, effectively unique, and parameter-free: ‘A dark matter halo is comprised of the collection of particles orbiting in their own self-generated potential’. This definition is enabled by the fact that, even with as few as ≈300 particles per halo, nearly every particle in the vicinity of a halo can be uniquely classified as either orbiting or infalling based on its dynamical history. For brevity, we refer to haloes selected in this way as physical haloes. We demonstrate that (1) the mass function of physical haloes is Press–Schechter, provided the critical threshold for collapse is allowed to vary slowly with peak height; and (2) the peak-background split prediction of the clustering amplitude of physical haloes is statistically consistent with the simulation data, with accuracy no worse than ≈5 per cent.

Astronomy & Astrophysics↗

Emulating ab initio computations of infinite nucleonic matter

We construct efficient emulators for the computation of the infinite nuclear matter equation of state. These emulators are based on the subspace-projected coupled-cluster method for which we here develop a new algorithm called small-batch voting to eliminate spurious states that might appear when emulating quantum many-body methods based on a non-Hermitian Hamiltonian. The efficiency and accuracy of these emulators facilitate a rigorous statistical analysis within which we explore nuclear matter predictions for > 10 6 different parametrizations of a chiral interaction model with explicit Δ -isobars at next-to-next-to leading order. Constrained by nucleon-nucleon scattering phase shifts and bound-state observables of light nuclei up to He 4 , we use history matching to identify nonimplausible domains for the low-energy coupling constants of the chiral interaction. Within these domains we perform a Bayesian analysis using sampling and importance resampling with different likelihood calibrations and study correlations between interaction parameters, calibration observables in light nuclei, and nuclear matter saturation properties. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Verification Testing of OLI Systems Mixed Solvent Electrolyte Model for the Na-K-Mg-Ca-H-Cl-SO 4 -OH-HCO 3 -CO 3 -CO 2 -H 2 ) System to High Ionic Strength at 25°C

This technical report summarizes model verification results and summary statistics for 41 evaporite mineral solubility cases evaluated by Savannah River National Laboratory using OLI Systems’ aqueous electrolyte thermodynamic modeling software. The 41 verification cases containing a total of 60 solubility curves comprise mineral solubility data from low to high ionic strength at 25°C for the eight-component system Na-K-Mg-Ca-H-Cl-SO 4 -OH-HCO 3 -CO 3 -CO 2 -H 2 O as reported by Harvie et al. (1984). Thermodynamic calculations were executed using OLI Systems’ Stream Analyzer computation module within the OLI Studio software platform (Ver. 11.0, Rev. 11.0.1.9). The Mixed Solvent Electrolyte (MSE) thermodynamic framework was chosen for this investigation because of its superiority in modeling high ionic-strength inorganic salt solutions and actinide redox chemistry and solubility, both of which are relevant to the geological repository conditions at the Waste Isolation Pilot Plant in Carlsbad, New Mexico. Mineral solubility data in various inorganic salt solutions were digitized and extracted from figures generated by Harvie et al. (1984). For each of the 60 solubility curves, a case-specific chemistry model and input file were generated in OLI Studio using OLI Stream Analyzer and the MSE (H 3 O + ion) public databank provided by OLI Systems. Model simulation results were exported to Microsoft Excel to calculate summary statistics and to generate graphs comparing the OLI model predictions to the solubility data. Summary statistics include residuals (model – data) and concordance (accuracy × precision, where precision is indicated by the Pearson correlation coefficient and accuracy accounts for bias and scale differential). Private databanks were not developed, and activity coefficient model regressions were not performed to improve OLI model fits to the data. Of the 41 model verification plots, 83% have a mean of the percent residuals less than or equal to 25%. Similarly, 75% display a concordance greater than or equal to 0.75. Only seven of the 41 verification plots fail to show good agreement between the model and data. Of these seven, three are relevant to the WIPP repository because they involve the Mg-OH-Cl-SO 4 -CO 3 aqueous system. The remaining four address salt solubilities at the pH extremes (strong acid and strong base). It should be noted that in two of the three Mg-OH-Cl-SO 4 -CO 3 system cases, the regressed Harvie et al. (1984) solubility curve also deviated from the data. Lack of agreement between the OLI model-predicted solubility curves and the data is attributable to one or more of the following: specific solid species are not included in the OLI MSE databank; there is significant variation among the different solubility datasets chosen by Harvie et al. (1984); the OLI MSE model’s thermodynamic parameters were determined using different solubility datasets; and the activity coefficient parameters for certain relevant ion-ion and ion-molecule pairs have not been optimized via data regression. Two recommendations for future work are to (1) evaluate solubility data for the Mg-OH-Cl-SO 4 -CO 3 system at high ionic strength and, if necessary, develop a private OLI MSE database that includes missing species and, where necessary, regressed standard state properties and interaction parameters; (2) perform similar verification testing of the OLI model for actinide solubility data.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗