Cosmology from HSC Y1 weak lensing data with combined higher-order statistics and simulation-based inference
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Understanding the collective behavior of complex spin textures, such as lattices of magnetic skyrmions, is of fundamental importance for exploring and controlling the emergent ordering of these spin textures and inducing phase transitions. It is also critical to understand the skyrmion–skyrmion interactions for applications such as magnetic skyrmion-enabled reservoir or neuromorphic computing. Magnetic skyrmion lattices can be studied using in situ Lorentz transmission electron microscopy (LTEM), but quantitative and statistically robust analysis of the skyrmion lattices from LTEM images can be difficult. In this work, we show that a convolutional neural network, trained on simulated data, can be applied to perform segmentation of spin textures and to extract quantitative data, such as spin texture size and location, from experimental LTEM images, which cannot be obtained manually. This includes quantitative information about skyrmion size, position, and shape, which can, in turn, be used to calculate skyrmion–skyrmion interactions and lattice ordering. We apply this approach to segmenting images of Néel skyrmion lattices so that we can accurately identify skyrmion size and deformation in both dense and sparse lattices. The model is trained using a large set of micromagnetic simulations as well as simulated LTEM images. This entirely open-source training pipeline can be applied to a wide variety of magnetic features and materials, enabling large-scale statistical studies of spin textures using LTEM.
Shortly after its discovery, General Relativity (GR) was applied to predict the behavior of our Universe on the largest scales, and later became the foundation of modern cosmology. Its validity has been verified on a range of scales and environments from the Solar system to merging black holes. However, experimental confirmations of GR on cosmological scales have so far lacked the accuracy one would hope for — its applications on those scales being largely based on extrapolation and its validity there sometimes questioned in the shadow of the discovery of the unexpected cosmic acceleration. Future astronomical instruments surveying the distribution and evolution of galaxies over substantial portions of the observable Universe, such as the Dark Energy Spectroscopic Instrument (DESI), will be able to measure the fingerprints of gravity and their statistical power will allow strong constraints on alternatives to GR. In this paper, based on a set of N-body simulations and mock galaxy catalogs, we study the predictions of a number of traditional and novel summary statistics beyond linear redshift distortions in two well-studied modified gravity models — chameleon f(R) gravity and a braneworld model — and the potential of testing these deviations from GR using DESI. These summary statistics employ a wide array of statistical properties of the galaxy and the underlying dark matter field, including two-point and higher-order statistics, environmental dependence, redshift space distortions and weak lensing. We find that they hold promising power for testing GR to unprecedented precision. The major future challenge is to make realistic, simulation-based mock galaxy catalogs for both GR and alternative models to fully exploit the statistic power of the DESI survey (by matching the volumes and galaxy number densities of the mocks to those in the real survey) and to better understand the impact of key systematic effects. Using these, we identify future simulation and analysis needs for gravity tests using DESI.
Trapped atomic ions find wide applications ranging from precision measurement to quantum information science and quantum computing. Beryllium ions are widely used due to the light mass and convenient atomic structure of beryllium; however, conventional ion loading from thermal ovens exerts undesirable gas loads for a prolonged duration. Here, we demonstrate a method to rapidly produce pure linear chains of beryllium ions with pulsed laser ablation, serving as a starting point for large-scale quantum information processing. Our method is fast compared to thermal ovens, reduces the gas load to only 10 -12 Torr (10 -10 Pa) level, yields a short recovery time of a few seconds, and also eliminates the need for a deep ultraviolet laser for photoionization. We also study the loading dynamics, which show non-Poissonian statistics in the presence of sympathetic cooling. In addition, we apply feedback control to obtain defect-free ion chains with desirable lengths.
In this article, we perform energy-dependent calculations of independent and cumulative fission product yields for 235 U, 238 U, and 239 Pu in the first chance fission region. Starting with the primary fission fragment distributions taken from available experimental data and analytical functions based on assumptions for the excitation energy and spin-parity distributions, the Hauser-Feshbach statistical decay treatment for fission fragment de-excitation is applied to more than 1,000 fission fragments for the incident neutron energies up to 5 MeV. The calculated independent yields are then used as an input of β-decay calculations to produce the cumulative yield, and summation calculations are performed. Model parameters in these procedures are adjusted by applying the Bayesian technique at the thermal energy for 235 U and 239 Pu and in the fast energy range for 238 U. The calculated fission observable quantities, such as the energy-dependent cumulative yields, and prompt and delayed neutron yields, are compared with available experimental data. We also study the possible impact of the second chance fission opening on the energy dependence of the delayed neutron yield by on the opening of second chance fission on the energy dependence calculation.
Abstract Measures of inconsistency and tension between datasets have become an essential part of cosmological analyses. It is important to accurately evaluate the significance of such tensions when present. We propose here a Bayesian interpretation of inconsistency measures that can extract information about physical inconsistencies in the presence of data scatter. This new framework is based on the conditional probability distribution of the level of physical inconsistency given the obtained value of the measure. We use the index of inconsistency as a case study to illustrate the new interpretation framework, but this can be generalized to other metrics. Importantly, there are two aspects in the quantification of inconsistency that behave differently as the number of model parameters increases. The first is the probability for the level of physical inconsistency to reach a threshold which drops with the increase of the number of parameters under consideration. The second is the actual level of physical inconsistency which remains rather insensitive to such an increase in parameters. The difference between these two aspects is often overlooked, which leads to a long-standing ambiguity: when a given inconsistency is found between two constraints, its “significance” seems to be lower when considered in a higher-dimensional parameter space. This ambiguity is resolved by the Bayesian interpretation we introduce in this work because the conditional probability distribution includes all the statistical information of the level of physical inconsistency. Finally, we apply the Bayesian interpretation to examine the (in)consistency between Planck versus the Cepheid-based local measurement, the Dark Energy Survey (DES), the Atacama Cosmology Telescope (ACT) and WMAP. We confirm and revisit the degrees of previous physical inconsistencies and show the stability of the new interpretation with respect to the number of cosmological parameters compared to the commonly used n-σ interpretation when applied to cosmological tensions in multi-parameter spaces.
An accurate calibration of the source redshift distribution p(z) is a key aspect in the analysis of cosmic shear data. This, one way or another, requires the use of spectroscopic or high-quality photometric samples. However, the difficulty to obtain colour-complete spectroscopic samples matching the depth of weak lensing catalogs means that the analyses of different cosmic shear datasets often use the same samples for redshift calibration. This introduces a source of statistical and systematic uncertainty that is highly correlated across different weak lensing datasets, and which must be accurately characterised and propagated in order to obtain robust cosmological constraints from their combination. In this paper we introduce a method to quantify and propagate the uncertainties on the source redshift distribution in two different surveys sharing the same calibrating sample. The method is based on an approximate analytical marginalisation of the p(z) statistical uncertainties and the correlated marginalisation of residual systematics. We apply this method to the combined analysis of cosmic shear data from the DESY1 data release and the HSC-DR1 data, using the COSMOS 30-band catalog as a common redshift calibration sample. We find that, although there is significant correlation in the uncertainties on the redshift distributions of both samples, this does not change the final constraints on cosmological parameters significantly. The same is true also for the impact of residual systematic uncertainties from the errors in the COSMOS 30-band photometric redshifts. Additionally, we show that these effects will still be negligible in Stage-IV datasets. Finally, the combination of DESY1 and HSC-DR1 allows us to constrain the "clumpiness" parameter to S 8 = ${0.768}_{-0.017}^{+0.021}$. This corresponds to a ~√(2) improvement in uncertainties with respect to either DES or HSC alone.
We present an extension to a Sunyaev–Zel’dovich Effect (SZE) selected cluster catalogue based on observations from the South Pole Telescope (SPT); this catalogue extends to lower signal to noise than the previous SPT–SZ catalogue and therefore includes lower mass clusters. Optically derived redshifts, centres, richnesses, and morphological parameters together with catalogue contamination and completeness statistics are extracted using the multicomponent matched filter (MCMF) algorithm applied to the S/N > 4 SPT–SZ candidate list and the Dark Energy Survey (DES) photometric galaxy catalogue. The main catalogue contains 811 sources above S/N = 4, has 91 per cent purity, and is 95 per cent complete with respect to the original SZE selection. It contains in total 50 per cent more clusters and twice as many clusters above z = 0.8 in comparison to the original SPT-SZ sample. The MCMF algorithm allows us to define subsamples of the desired purity with traceable impact on catalogue completeness. As an example, we provide two subsamples with S/N > 4.25 and S/N > 4.5 for which the sample contamination and cleaning-induced incompleteness are both as low as the expected Poisson noise for samples of their size. The subsample with S/N > 4.5 has 98 per cent purity and 96 per cent completeness and is part of our new combined SPT cluster and DES weak-lensing cosmological analysis. We measure the number of false detections in the SPT-SZ candidate list as function of S/N, finding that it follows that expected from assuming Gaussian noise, but with a lower amplitude compared to previous estimates from simulations.
Host-associated microbial communities are shaped by extrinsic and intrinsic factors to the holobiont organism. Environmental factors and microbe-microbe interactions act simultaneously on the microbial community structure, making the microbiome dynamics challenging to predict. The coral microbiome is essential to the health of coral reefs and sensitive to environmental changes. Here, we develop a dynamic model to determine the microbial community structure associated with the surface mucus layer (SML) of corals using temperature as an extrinsic factor and microbial network as an intrinsic factor. The model was validated by comparing the predicted relative abundances of microbial taxa to the relative abundances of microbial taxa from the sample data. The SML microbiome from Pseudodiploria strigosa was collected across reef zones in Bermuda, where inner and outer reefs are exposed to distinct thermal profiles. A shotgun metagenomics approach was used to describe the taxonomic composition and the microbial network of the coral SML microbiome. By simulating the annual temperature fluctuations at each reef zone, the model output is statistically identical to the observed data. The model was further applied to six scenarios that combined different profiles of temperature and microbial network to investigate the influence of each of these two factors on the model accuracy. The SML microbiome was best predicted by model scenarios with the temperature profile that was closest to the local thermal environment, regardless of the microbial network profile. Our model shows that the SML microbiome of P. strigosa in Bermuda is primarily structured by seasonal fluctuations in temperature at a reef scale, while the microbial network is a secondary driver. Coral microbiome dysbiosis (i.e., shifts in the microbial community structure or complete loss of microbial symbionts) caused by environmental changes is a key player in the decline of coral health worldwide. Multiple factors in the water column and the surrounding biological community influence the dynamics of the coral microbiome. However, by including only temperature as an external factor, our model proved to be successful in describing the microbial community associated with the surface mucus layer (SML) of the coral P. strigosa. The dynamic model developed and validated in this study is a potential tool to predict the coral microbiome under different temperature conditions.
Background: To discover pharmacotherapy prescription patterns and their statistical associations with outcomes through a clinical pathway inference framework applied to real-world data. Methods: We apply machine learning steps in our framework using a 2006 to 2020 cohort of veterans with major depressive disorder (MDD). Outpatient antidepressant pharmacy fills, dispensed inpatient antidepressant medications, emergency department visits, self-harm, and all-cause mortality data were extracted from the Department of Veterans Affairs Corporate Data Warehouse. Results: Our MDD cohort consisted of 252,179 individuals. During the study period there were 98,417 emergency department visits, 1,016 cases of self-harm, and 1,507 deaths from all causes. The top ten prescription patterns accounted for 69.3% of the data for individuals starting antidepressants at the fluoxetine equivalent of 20-39 mg. Additionally, we found associations between outcomes and dosage change. Conclusions: For 252,179 Veterans who served in Iraq and Afghanistan with subsequent MDD noted in their electronic medical records, we documented and described the major pharmacotherapy prescription patterns implemented by Veterans Health Administration providers. Ten patterns accounted for almost 70% of the data. Associations between antidepressant usage and outcomes in observational data may be confounded. The low numbers of adverse events, especially those associated with all-cause mortality, make our calculations imprecise. Furthermore, our outcomes are also indications for both disease and treatment. Despite these limitations, we demonstrate the usefulness of our framework in providing operational insight into clinical practice, and our results underscore the need for increased monitoring during critical points of treatment.
This technical final report submitted to DOE/NETL presents all the research activities performed during the entirety of DE-FE0031660 project-Emissions Mitigation Technology for Advanced Water-Lean Solvent-Based CO 2 Capture Processes which spans from October 2018 through March 2022. RTI International has been conducting studies from fundamental and operational aspects to reduce the overall amine emissions from the advanced Water-Lean Solvent (WLS) systems, specifically RTI’s Non-Aqueous Solvent (NAS). This technical final report will highlight the key findings from project which align closely to the project objectives which are: Identify the contribution of vapor loss, entrainment, and aerosols to the overall emissions of water-lean systems; Determine the significance of CO 2 capture system operating parameters to the amine emissions; Develop an emissions model based on critical operating parameters; Evaluate the effectiveness of emissions mitigation devices to reduce the amine emissions to <1 ppm under flue coal-fired flue gas; and, Determine the contribution of the ECTs to the overall CO 2 capture cost. The following are the key findings based on numerous tests using both lab-scale setups and parametric testing performed at RTI’s Bench-scale Gas Absorption System (BsGAS). During the BP1, the aerosol generation system and monitoring equipment were installed at BsGAS to produce and determine the aerosol characteristics during the NAS CO 2 capture process. The aerosol produced by this setup produced aerosols with the peak diameter and concentration of 50 micron and 1.2E10 7 cm -3 , respectively. These particle sizes and concentrations are matched to those observed in the actual coal-fired power plant flue gases and expected to be found at the absorber inlet of the CO 2 capture system. Over 1,300 hours of parametric testing have been conducted to evaluate the impact of the aerosols and operating conditions during the CO 2 capture with NAS on the overall amine emissions in the treated flue gas. At the worse condition tested, the presence of the aerosols in the flue gas could increase the overall emissions by 10X compared to the baseline emissions from NAS’s vapor pressure. CO 2 capture rate was found to be a main factor impacting the overall emissions as well as aerosol size and concentrations in the absorber off-gas. The higher CO 2 capture rate, the higher amine emissions in the treated gas. The temperature difference between the temperature bulge seen in the absorber and the water wash temperature also impacts the particle growth where the larger the temperature difference, the more amine emissions from aerosols in the treated gas. The majority of the aerosols did not grow substantially in the system, and the particle concentrations remained nearly constant between the absorber inlet and wash outlet. Only a small portion of the particles were found to grow significantly. The high efficiency demister with mesh size of 5-10 micron can be installed to remove a portion of the aerosols from the gas stream leaving the water wash. Overall, these results from parametric testing have established the emission baseline and validate our assumption on the need of emission control technologies (ECT) in order to minimize the emissions from the baseline NAS CO 2 capture process. Over 2,000 of BsGAS operating hours was used to investigate a handful of process improvements which led to a selection of the vital few changes that effectively control the amine emissions. These process improvements are lime-coated-filters for absorber gas inlet, advanced demister at the top of the absorber, a second water wash with amine recovery unit were designed, installed, and tested at BsGAS at the end of BP1. The result showed that the NAS CO 2 capture process with these additional emission control devices could lower the amine emission in the treated gas to about 1 ppm using a simulated coal-fire flue gas stream. The main contributor in lowering the amine emission came from the second water wash with amine recovery unit where the amine concentration in the scrubbing water was kept below 2 wt% through a continuous amine removal via an adsorbent bed, resulting in a low amine vapor pressure. The adsorbent bed was regenerated via a direct steam regeneration and the recover amine was returned to the absorber to minimize wastewater and makeup amine. A flue gas generation system was designed and installed during the first half of BP2 to support the emission testing using a real coal-derived flue gas. The system is capable of generating both coal- and natural gas- derived flue gases with the composition of the gaseous species highly resemble to that of the power plant flue gases. The particulates detected in the coal-derived flue gas showed the mean diameter of 1 micron. The CO 2 capture operating was then proceed using the real coal-derived flue gas where the amine emission was controlled to be about 0-3 ppm for the total run time of about 200 hours. Similar testing was conducted with natural gas-derived flue gas and the result showed a highly amine emission of 30 ppm under the total run time of 200 hours. The Principal Component Analysis (PCA) and the Partial Least Squares Projection to Latent Structures (PLS) techniques were applied to the parametric testing data to derive a multivariate statistical model. The model was validated and trained with half of the data collected, and the predictive ability of the model was evaluated using the remaining half of the data. The resulting empirical model was capable of predicting the overall emissions from the NAS process without the ECTs with ±15% accuracy (average absolute deviation, AAD) in BP1. As more emission data were obtained under the real coal-flue gas in the BP2, the model incorporated these new set of data to reflect the final process configuration, operating parameters, and amine emission. This results in the updated empirical model predicting the amine emission from the NAS CO 2 capture process with 84% goodness-of-fit (R 2 ), 85% predictability (Q 2 ), and 15% AAD. The study evaluates the use of RTI’s Non-Aqueous Solvent technology for 90% CO 2 capture from a net 650 MWe pulverized coal power plant, downstream of the flue-gas desulfurization unit. The captured CO 2 has a purity of > 95% CO 2 , and is dried, compressed to 15.3 MPa (2,215 psia), ready for sequestration. The analysis uses Case B12B from the DOE Baseline study on Bituminous Coal, Revision 4 where the Cansolv CO 2 capture plant is replaced by the RTI CO 2 Capture plant. The CO 2 capture plant has been sized to capture >90% CO 2 from flue gas derived from a net 650 MWe supercritical pulverized coal power plant. The CO 2 capture plant is equipped with emission control technologies that limits the amine emissions to < 1 ppm. Two different cases were evaluated for the technoeconomic study. The key difference between the two cases is the regenerator pressure. In Case 1, the regenerator operates at 0.195 MPa (28.3 psia), whereas in Case 2, the regenerator pressure is 0.44 MPa (64 psia) thus removing the need for the first stage of compression of the eight-stage compression train. Results from the TEA are compared against the DOE reference cases for SCPC plant with and without CO 2 Capture (Case B12A and Case B12B of the DOE Baseline study, respectively). Case 2 with CO 2 regeneration at higher pressure results in the lower cost of CO 2 capture. The total capital cost of the capture process has been estimated using 2018 dollars in Aspen Process Economic Analyzer and was estimated to be $579 MM. The capture plant operation leads to a total parasitic power loss rate of 96 MWe, resulting in a decrease in pulverized coal power plant efficiency of 7.8% points. The resulting cost of electric power increases from 64.4 mills/kWh, for no capture, to 97.5 mills/kWh, with 90% capture, an increase of 51% in the COE. The cost of capturing 90% CO 2 was estimated to be $38.2/tonne-CO 2 , and meets the DOE target of $40/t-CO 2 . Emission control technologies (ECT) investigated in this project includes a second water wash with use of activated carbon beds for removal of amine from the wash water prior to recirculation in the water wash. These ECT allow operation of the CO 2 capture plant with < 1 ppm amine emissions with the treated flue gas and contributes to $2.4/t-CO 2 captured. Amine emissions derived from thermal and oxidative degradations were investigated under this project along with the emissions derived from aerosols for the NAS system. The thermally degraded of the lean NAS showed less than 4% decreased of the original total amine content in the NAS at 150 °C while the result obtained at 120 °C showed no drop in total amine content, suggesting that thermal degradation of the NAS is minimal. These results also suggested that the thermally degraded species are not likely formed and contributed to the emissions due to the low regeneration temperature of the NAS at 90-105 °C. The oxidative degradation, on the other hand, could become problematic as some of these oxidative degraded species were observed during the NAS-5 testing at National Carbon Capture Center (NCCC) and SINTEF in our previous project. The rapid screening of selected inhibitors suggested that oxidative degradation of NAS can be suppressed using thiol containing compounds in amounts of at least 1 mol%. The detailed mechanistic degradation pathway was conceived for a specific amine used in NAS formulation during BP2. he reduction of the nitrosamines caused by the NO x present in the flue gas was also examined. The study suggested that the thermo-chemical treatment of the NAS solvent would be a more effective and economically viable compared to removing NO x at the DCC.
This work presents a novel method for generating electricity price scenarios from statistical properties of past electricity prices using a hybrid statistical and reduced-form stochastic model. Previous work in applying stochastic differential equations (SDE) to model electricity prices has focused on daily average prices. To extend stochastic price generation methods to hourly or sub-hourly pricing, we address several weaknesses in the state-of-the-art: (1) we replace the mean-reversion component of the SDE with an ARIMA process that is better able to characterize the daily and weekly trends; (2) we extend the price-spike, or jump process to account for conditional probabilities of price spikes occurring in consecutive time steps by replacing the traditional Poisson process for modeling jumps with a generalized point process model inspired by brain neuron models; and (3) we replace the traditional method of estimating spike intensity with empirical variance with a Markov process based on observed price spike intensity transitions. The method is demonstrated with electricity prices from the US ERCOT market and a use-case example is provided for bidding an energy storage unit into the day-ahead and real-time energy markets of ERCOT using stochastic optimization methods. Results show that the the synthetic price model out performs a (naive) persistence forecast model by resulting in 24% to 47% more in profits over 168 simulated days.
Semi-arid ecosystems, like those in the American Southwest, exert a massive impact on the interannual variability of carbon and water cycling. Unfortunately, these carbon and water fluxes are notoriously difficult to predict due to their high spatial and temporal variability, which is poorly captured by the current generation of vegetation models. Indeed, this region is exemplified by the ‘hot spots and hot moments’ concept, which states that small areas in space (‘hot spots’) and transient moments in time (‘hot moments’) exert an outsized influence on biogeochemical cycling. However, the factors that regulate these pulses in biogeochemical activity are unknown, as is their variability across space and time. These uncertainties severely limit efforts to better represent hot spots and hot moments in models. Here, we seek to develop a generalized method for detecting and quantifying the importance of hot spots and hot moments from individual plant to regional scales. Underpinning this method is our recently developed statistical approach for identifying hot spots and hot moments. By applying this method to semi-continuous measurements of plant water status, a depth profile of soil water potential, and ecosystem fluxes via eddy covariance, we will track the fate of water through the soil-plant-atmosphere continuum and identify the mechanistic drivers of these transient pulses in biogeochemical activity. Then, we will expand this approach across a broad network of Ameriflux towers, and apply a machine learning approach that will allow us to upscale measurements of hot spots and hot moments across the American Southwest and quantify their impact on carbon and water cycles. These products will allow us to identify hot spots and hot moments across spatio-temporal scales and will serve as crucial data sources for validating a new generation of models that can better capture highly dynamic carbon and water fluxes. The proposed method will be easily transferable across biomes and will serve as a framework for future research on hot spots and hot moments across the plant ecophysiology, biometeorology, and vegetation modeling communities.
Deep learning (DL) models have enjoyed increased attention in recent years because of their powerful predictive capabilities. While many successes have been achieved, standard deep learning methods suffer from a lack of uncertainty quantification (UQ). While the development of methods for producing UQ from DL models is an active area of current research, little attention has been given to the quality of the UQ produced by such methods. In order to deploy DL models to high-consequence applications, high-quality UQ is necessary. This report details the research and development conducted as part of a Laboratory Directed Research and Development (LDRD) project at Sandia National Laboratories. The focus of this project is to develop a framework of methods and metrics for the principled assessment of UQ quality in DL models. This report presents an overview of UQ quality assessment in traditional statistical modeling and describes why this approach is difficult to apply in DL contexts. An assessment on relatively simple simulated data is presented to demonstrate that UQ quality can differ greatly between DL models trained on the same data. A method for simulating image data that can then be used for UQ quality assessment is described. A general method for simulating realistic data for the purpose of assessing a model’s UQ quality is also presented. A Bayesian uncertainty framework for understanding uncertainty and existing metrics is described. Research that came out of collaborations with two university partners are discussed along with a software toolkit that is currently being developed to implement the UQ quality assessment framework as well as serve as a general guide to incorporating UQ into DL applications.
This report demonstrates that applying graph theory techniques provides a way to obtain sufficient statistics in finding errors when testing complex state machines. It discusses how to define the tests, then demonstrates how to automatically generate test suites that diversify test cases, subject to constraints. If included within a continuous integration approach, these constructs provide an unbiased means to systematically check for errors within the latest controller software release.
This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.
Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.
We present a novel technique to incorporate precision calculations from quantum chromodynamics into fully differential particle-level Monte Carlo simulations. By minimizing an information-theoretic quantity subject to constraints, our reweighted Monte Carlo incorporates systematic uncertainties absent in individual Monte Carlo predictions, achieving consistency with the theory input in precision and its estimated systematic uncertainties. Our method can be applied to arbitrary observables known from precision calculations, including multiple observables simultaneously. It generates strictly positive weights, thus offering a clear path to statistically powerful and theoretically precise computations for current and future collider experiments. As a proof of concept, we apply our technique to event-shape observables at electron-positron colliders, leveraging existing precision calculations of thrust. Our analysis highlights the importance of logarithmic moments of event shapes, which have not been previously studied in the collider physics literature.