Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Characterization of contaminants in the Lyman-alpha forest auto-correlation with DESI

Baryon Acoustic Oscillations can be measured with sub-percent precision above redshift two with the Lyman-α (Lyα) forest auto-correlation and its cross-correlation with quasar positions. This is one of the key goals of the Dark Energy Spectroscopic Instrument (DESI) which started its main survey in May 2021. We present in this paper a study of the contaminants to the Lyα forest which are mainly caused by correlated signals introduced by the spectroscopic data processing pipeline as well as astrophysical contaminants due to foreground absorption in the intergalactic medium. Notably, an excess signal caused by the sky background subtraction noise is present in the Lyα auto-correlation in the first line-of-sight separation bin. We use synthetic data to isolate this contribution, we also characterize the effect of spectro-photometric calibration noise, and propose a simple model to account for both effects in the analysis of the Lyα forest. We then measure the auto-correlation of the quasar flux transmission fraction of low redshift quasars, where there is no Lyα forest absorption but only its contaminants. We demonstrate that we can interpret the data with a two-component model: data processing noise and triply ionized Silicon and Carbon auto-correlations. This result can be used to improve the modeling of the Lyα auto-correlation function measured with DESI.

79 ASTRONOMY AND ASTROPHYSICS↗

MOSAIC: a joint modeling methodology for combined circadian and non-circadian analysis of multi-omics data

Abstract Motivation Circadian rhythms are approximately 24-h endogenous cycles that control many biological functions. To identify these rhythms, biological samples are taken over circadian time and analyzed using a single omics type, such as transcriptomics or proteomics. By comparing data from these single omics approaches, it has been shown that transcriptional rhythms are not necessarily conserved at the protein level, implying extensive circadian post-transcriptional regulation. However, as proteomics methods are known to be noisier than transcriptomic methods, this suggests that previously identified arrhythmic proteins with rhythmic transcripts could have been missed due to noise and may not be due to post-transcriptional regulation. Results To determine if one can use information from less-noisy transcriptomic data to inform rhythms in more-noisy proteomic data, and thus more accurately identify rhythms in the proteome, we have created the Multi-Omics Selection with Amplitude Independent Criteria (MOSAIC) application. MOSAIC combines model selection and joint modeling of multiple omics types to recover significant circadian and non-circadian trends. Using both synthetic data and proteomic data from Neurospora crassa, we showed that MOSAIC accurately recovers circadian rhythms at higher rates in not only the proteome but the transcriptome as well, outperforming existing methods for rhythm identification. In addition, by quantifying non-circadian trends in addition to circadian trends in data, our methodology allowed for the recognition of the diversity of circadian regulation as compared to non-circadian regulation. Availability and implementation MOSAIC’s full interface is available at https://github.com/delosh653/MOSAIC. An R package for this functionality, mosaic.find, can be downloaded at https://CRAN.R-project.org/package=mosaic.find. Supplementary information Supplementary data are available at Bioinformatics online.

De los Santos, Hannah↗

Robust and optimal alignment of high-dimensional data using maximum likelihood estimation through a random sample consensus framework

Abstract Correcting spatial orientations of groups of high-dimensional data sets such that they are all in a consistent coordinate system is often a time-consuming and error-prone process. Automation of this process can be accomplished by using Generalized Procrustes Analysis to estimate the relative orientations among a population of high-dimensional data sets. A least squares Procrustes solution is applied through a maximum likelihood estimation and random sample consensus framework for robustness. The likelihood model is comprised of a mixture distribution where inliers are modeled using t -distribution and outliers from a uniform distribution. Applications will focus on a synthetic data set that emulates triaxial acceleration data and also real shock data from a population of triaxial accelerometers. Outliers represent either non-rigid body responses, environmental noise, and/or sensor and data acquisition issues. The intended application for the methodology is to robustly automate the rotation of populations of experimentally collected triaxial accelerometer data sets to a single global coordinate system.

LOSAC↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

A Benchmark to Test Generalization Capabilities of Deep Learning Methods to Classify Severe Convective Storms in a Changing Climate

Abstract This is a test case study assessing the ability of deep learning methods to generalize to a future climate (end of 21st century) when trained to classify thunderstorms in model output representative of the present‐day climate. A convolutional neural network (CNN) was trained to classify strongly rotating thunderstorms from a current climate created using the Weather Research and Forecasting model at high‐resolution, then evaluated against thunderstorms from a future climate and found to perform with skill and comparatively in both climates. Despite training with labels derived from a threshold value of a severe thunderstorm diagnostic (updraft helicity), which was not used as an input attribute, the CNN learned physical characteristics of organized convection and environments that are not captured by the diagnostic heuristic. Physical features were not prescribed but rather learned from the data, such as the importance of dry air at mid‐levels for intense thunderstorm development when low‐level moisture is present (i.e., convective available potential energy). Explanation techniques also revealed that thunderstorms classified as strongly rotating are associated with learned rotation signatures. Results show that the creation of synthetic data with ground truth is a viable alternative to human‐labeled data and that a CNN is able to generalize a target using learned features that would be difficult to encode due to spatial complexity. Most importantly, results from this study show that deep learning is capable of generalizing to future climate extremes and can exhibit out‐of‐sample robustness with hyperparameter tuning in certain applications.

54 ENVIRONMENTAL SCIENCES↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Testing Fractal Methods on Observed and Simulated Solar Magnetograms

The term "magnetic complexity" has not been sufficiently quantified. To accomplish this, we must understand the relationship between the observed magnetic field of solar active regions and fractal dimension measurements. Using data from the Marshall Space Flight Center's vector magnetograph ranging from December 1991 to July 2001, we compare the results of several methods of calculating a fractal dimension, e.g., Hurst coefficient, the Higuchi method, power spectrum, and 2-D Wavelet Packet Analysis. In addition, we apply these methods to synthetic data, beginning with representations of very simple dipole regions, ending with regions that are magnetically complex.

Adams, M.↗

Revisiting the Kinematics of the Cylinder Test

The cylinder expansion experiment is a well-established performance test for condensed explosives that is utilized routinely to determine pressure-energy-volume relationships for detonation products. The modern cylinder test employs optical interferometric techniques to measure velocity, as opposed to older realizations where streak cameras were commonplace. Despite their widespread use, questions sometimes remain as to what kind of data the velocity diagnostics in a cylinder test provide along with how to interpret such measurements and relate them to physical quantities of interest. Here in this study, equations are derived that fully describe the kinematics of the cylinder wall during expansion. These equations are then applied to experimental as well as synthetic data generated via a hydrodynamic simulation in order to verify the mathematical framework developed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Limited Angle Reconstruction Method for Reconstructing Terrestrial Plasmaspheric Densities from EUV Images

A new method for reconstructing the global 3D distribution of plasma densities in the plasmasphere from a limited number of 2D views is presented. The method is aimed at using data from the Extreme Ultra Violet (EUV) sensor on NASA s Imager for Magnetopause-to-Aurora Global Exploration (IMAGE) satellite. Physical properties of the plasmasphere are exploited by the method to reduce the level of inaccuracy imposed by the limited number of views. The utility of the method is demonstrated on synthetic data.

Newman, Timothy↗

Chemical Modeling of the Reactivity of Short-Lived Greenhouse Gases: A Model Inter-Comparison Prescribing a Well-Measured, Remote Troposphere

We develop a new protocol for merging in situ measurements with 3-D model simulations of atmospheric chemistry with the goal of integrating over the data to identify the most reactive air parcels in terms of tropospheric production and loss of the greenhouse gases ozone and methane. Presupposing that we can accurately measure atmospheric composition, we examine whether models constrained by such measurements agree on the chemical budgets for ozone and methane. In applying our technique to a synthetic data stream of 14,880 parcels along 180W, we are able to isolate the performance of the photochemical modules operating within their global chemistry-climate and chemistry-transport models, removing the effects of modules controlling tracer transport, emissions, and scavenging. Differences in reactivity across models are driven only by the chemical mechanism and the diurnal cycle of photolysis rates, which are driven in turn by temperature, water vapor, solar zenith angle, clouds, and possibly aerosols and overhead ozone, which are calculated in each model. We evaluate six global models and identify their differences and similarities in simulating the chemistry through a range of innovative diagnostics. All models agree that the more highly reactive parcels dominate the chemistry (e.g., the hottest 10% of parcels control 25-30% of the total reactivities), but do not fully agree on which parcels comprise the top 10%. Distinct differences in specific features occur, including the regions of maximum ozone production and methane loss, as well as in the relationship between photolysis and these reactivities. Unique, possibly aberrant, features are identified for each model, providing a benchmark for photochemical module development. Among the 6 models tested here, 3 are almost indistinguishable based on the inherent variability caused by clouds, and thus we identify 4, effectively distinct, chemical models. Based on this work, we suggest that water vapor differences in model simulations of past and future atmospheres may be a cause of the different evolution of tropospheric O3 and CH4, and lead to different chemistry-climate feedbacks across the models.

greenhouse gases↗

A generalized forward fit for neutron detectors with energy-dependent response functions

To date, most analysis of neutron time-of-flight data from inertial confinement fusion experiments has focused on the relatively small range of energies corresponding to the primary neutrons from DD and DT fusion, and have therefore employed instrument response functions (IRF’s) corresponding to monoenergetic 2.45-MeV or 14.03-MeV neutrons. For analysis of time-of-flight signals corresponding to broader ranges of neutron energies, accurate treatment of the data requires the use of an energy-dependent IRF. Here, this work describes interpolation of the IRF for neutrons of arbitrary energy, construction of an energy-dependent IRF, and application of this IRF in a forward fit via matrix multiplication. As an example of the application of this method, an analysis of synthetic data relevant to TT fusion experiments at the Omega Laser Facility is discussed. This example is used to illustrate the differences between a forward fit that uses an energy-dependent IRF and a forward fit that uses a monoenergetic IRF. Use of the energy-dependent IRF is shown to result in accurate inference of the fit parameters of interest.

47 OTHER INSTRUMENTATION↗

Magsat science investigations

Existing software is being modified to take any combination of component or scalar data in profile form and invert it to a discrete-source magnetization distribution for sources having arbitrary equal-area spacing. The option of constraining both source and directions and magnitude is included. Software for spectral depth-to-magnetic bottom estimates is under development. The software is to be thoroughly listed on synthetic data and applied to the NOO survey data and to NURE data for the southern Rio Grande Rift. Swanberg's silica geotemperature data for the U.S. was digitized for heat flow studies.

Source record↗

Optimized Profile Retrievals of Aerosol Microphysical Properties from Simulated Spaceborne Multiwavelength Lidar

This work is an expanded study of one previously published onretrievals of aerosol microphysical properties from space-borne multiwavelengthlidarmeasurements. The earlier studiesand this one weredone in the framework of the NASA Aerosol-Clouds-Ecosystems (now the Aerosol Clouds Convection and Precipitation) NASA mission. The focus here is on the capabilities of a simulated spacebornemultiwavelengthlidar system for retrieving aerosol complex refractive index (m = mr+ imi) and spectral single scattering albedo (SSA(λ)), although other bulk parameters such as effective (reff) radius and particle volume (V) and surface (S) concentrations are also studied. The novelty presented here is the use of recently published, case dependent optimized-constraints on the microphysical retrievals using three backscattering coefficients (β) at 355, 532 and 1064 nm and two extinction coefficients (α) at 355 and 532 nm, typically known as the stand-alone 3β+2α lidar inversion. Case-dependent optimized-constraints (CDOC) limit the ranges of refractive index, both real (mr) and imaginary (mi) parts, and of radii that are permitted in the retrievals. Such constraints are selected directly from the 3β+2α41measurements through an analysis of the relationship between spectral dependence of aerosol extinction-to-backscatter ratios (LR) and the Ångström exponent of extinction. The analyses presented here for different sets ofsize distributions and refractive indices reveal that the direct determination of CDOCareonly feasible for cases where the uncertaintiesin the input optical data areless than 15 %.Forthe same simulated spacebornesystem and yield than in Whiteman et al., (2018), we demonstrated that the use of CDOC as essential for the retrievals of refractive index and also largely improved retrieval of bulk parameters. A discussion of the global representativeness of CDOC is presented using simulated lidar data from a 24-hour satellite track using GEOS model output to initialize the lidar simulator.We found that CDOCare representative of many aerosol mixtures in spite of some outliers (e.g. highly hydrated particles) associatedwith the assumptions of bimodal size distributions and of the same refractive index for fine and coarse modes. Moreover, sensitivity tests performed using synthetic data reveal that retrievals of imaginary refractive index (mi) and SSA are extremely sensitive to β(355).

NASA Aerosol-Clouds-Ecosystems↗

Application of an energy-dependent instrument response function to analysis of nTOF data from cryogenic DT experiments

Neutron time-of-flight (nTOF) detectors are used to diagnose the conditions present in inertial confinement fusion (ICF) experiments and basic laboratory physics experiments performed on an ICF platform. The instrument response function (IRF) of these detectors is constructed by convolution of two components: an x-ray IRF and a neutron interaction response. The shape of the neutron interaction response varies with incident neutron energy, changing the shape of the total IRF. Analyses of nTOF data that span a broad range of energies must account for this energy-dependence in order to accurately infer plasma parameters and nuclear properties in ICF experiments. This work briefly reviews a matrix multiplication approach to convolution which allows for an energy-dependent change in the shape of the IRF. This method is applied to synthetic data resembling symmetric cryogenic DT implosions to examine the effect of the energy-dependent IRF on the inferred areal density. Here, results of forward fits that infer ion temperatures and areal densities from nTOF data collected during cryogenic DT experiments on OMEGA are also discussed.

47 OTHER INSTRUMENTATION↗

RADEMACHER COMPLEXITY REGULARIZATION FOR CORRELATION-BASED MULTIVIEW REPRESENTATION LEARNING

Deep correlation-based multiview representation learning techniques have become increasingly popular methods for extracting highly correlated representations from multiview data. However, their ability to find highly complex mappings between the views can also lead to overfitting and overly correlated representations. In this work, we propose a regularizer for this specific problem, based on the Rademacher complexity of the DNNs, tailored for multiview correlation maximization. We demonstrate that the proposed regularization leads to less noisy representations in synthetic data and improved performance of downstream tasks in real-world multiview datasets.

Kuschel, Maurice↗

MONTE CARLO CROSS SECTION LOOKUP KERNEL FOR THE CEREBRAS WSE-2 IN CSL

This is a small kernel that was used to collect data for an upcoming paper. We would like to have the code be open source so that the reviewers (and then readers) of the paper can see the whole code, and can reproduce/verify our results. This is not a fully featured application, it cannot produce any useful simulation results, it just executes a small abstracted kernel using synthetic data. The purpose of the kernel is to understand the basic performance characteristics of an HPC kernel on novel AI accelerator architectures. The main kernel is written in the CSL coding language for use with the Cerebras WSE-2 AI accelerator. The kernel represented is the Monte Carlo cross section lookup kernel, which is a small kernel used by the Monte Carlo neutral particle transport algorithm. There is also a baseline kernel written in CUDA that we will include in the repository to form a basis for comparing the WSE-2 to GPU.

Tramm, John↗

Dispersive and nondispersive 𝐾-matrix formalisms

The modeling of coupled-channel effects has become increasingly important due to the availability of highly precise data for a large variety of hadronic (re)scattering processes. The 𝐾-matrix is a powerful, yet comparatively simple, method to describe scattering amplitudes, including coupled-channel effects, with the aim of interpreting experimental data. Throughout the literature, a range of dispersive and nondispersive 𝐾-matrix methods are employed. Here, we compare the dispersive and nondispersive formulations in the context of the N/D method. It is shown that the methods are equivalent in the physical region under 𝐾-matrix reparametrization. Differences away from the physical region are examined. Applications to synthetic data are used to illustrate the effects of model choices concerning form factors and the application of dispersion relations, with the goal of clarifying best practices. We find no clear preference with regard to dispersive modeling. In contrast, we find that interpretational ambiguity of the bare model parameters—and even of the form of the bare model—is endemic, and recommend a thorough sampling of data and model spaces to assess conclusion robustness.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Metric to Quantify Shared Visual Attention in Two-Person Teams

Introduction: Critical tasks in high-risk environments are often performed by teams, the members of which must work together efficiently. In some situations, the team members may have to work together to solve a particular problem, while in others it may be better for them to divide the work into separate tasks that can be completed in parallel. We hypothesize that these two team strategies can be differentiated on the basis of shared visual attention, measured by gaze tracking. 2) Methods: Gaze recordings were obtained for two-person flight crews flying a high-fidelity simulator (Gontar, Hoermann, 2014). Gaze was categorized with respect to 12 areas of interest (AOIs). We used these data to construct time series of 12 dimensional vectors, with each vector component representing one of the AOIs. At each time step, each vector component was set to 0, except for the one corresponding to the currently fixated AOI, which was set to 1. This time series could then be averaged in time, with the averaging window time (t) as a variable parameter. For example, when we average with a t of one minute, each vector component represents the proportion of time that the corresponding AOI was fixated within the corresponding one minute interval. We then computed the Pearson product-moment correlation coefficient between the gaze proportion vectors for each of the two crew members, at each point in time, resulting in a signal representing the time-varying correlation between gaze behaviors. We determined criteria for concluding correlated gaze behavior using two methods: first, a permutation test was applied to the subjects' data. When one crew member's gaze proportion vector is correlated with a random time sample from the other crewmember's data, a distribution of correlation values is obtained that differs markedly from the distribution obtained from temporally aligned samples. In addition to validating that the gaze tracker was functioning reasonably well, this also allows us to compute probabilities of coordinated behavior for each value of the correlation. As an alternative, we also tabulated distributions of correlation coefficients for synthetic data sets, in which the behavior was modeled as a first-order Markov process, and compared correlation distributions for identical processes with those for disparate processes, allowing us to choose criteria and estimate error rates. 3) Discussion: Our method of gaze correlation is able to measure shared visual attention, and can distinguish between activities involving different instruments. We plan to analyze whether pilots strategies of sharing visual attention can predict performance. Possible measurements of performance include expert ratings from instructors, fuel consumption, total task time, and failure rate. While developed for two-person crews, our approach can be applied to larger groups, using intra-class correlation coefficients instead of the Pearson product-moment correlation.

Gontar, Patrick↗