Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Field size, length, and width distributions based on LACIE ground truth data

The development of agricultural remote sensing systems requires knowledge of agricultural field size distributions so that the sensors, sampling frames, image interpretation schemes, registration systems, and classification systems can be properly designed. Malila et al. (1976) studied the field size distribution for wheat and all other crops in two Kansas LACIE (Large Area Crop Inventory Experiment) intensive test sites using ground observations of the crops and measurements of their field areas based on current year rectified aerial photomaps. The field area and size distributions reported in the present investigation are derived from a representative subset of a stratified random sample of LACIE sample segments. In contrast to previous work, the obtained results indicate that most field-size distributions are not log-normally distributed. The most common field size observed in this study was 10 acres for most crops studied.

Pitts, D. E.↗

Detection of significant antiviral drug effects on COVID-19 with reasonable sample sizes in randomized controlled trials: A modeling study

Development of an effective antiviral drug for Coronavirus Disease 2019 (COVID-19) is a global health priority. Although several candidate drugs have been identified through in vitro and in vivo models, consistent and compelling evidence from clinical studies is limited. The lack of evidence from clinical trials may stem in part from the imperfect design of the trials. We investigated how clinical trials for antivirals need to be designed, especially focusing on the sample size in randomized controlled trials. A modeling study was conducted to help understand the reasons behind inconsistent clinical trial findings and to design better clinical trials. We first analyzed longitudinal viral load data for Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) without antiviral treatment by use of a within-host virus dynamics model. The fitted viral load was categorized into 3 different groups by a clustering approach. Comparison of the estimated parameters showed that the 3 distinct groups were characterized by different virus decay rates (p-value < 0.001). The mean decay rates were 1.17 d -1 (95% CI: 1.06 to 1.27 d -1 ), 0.777 d -1 (0.716 to 0.838 d -1 ), and 0.450 d -1 (0.378 to 0.522 d -1 ) for the 3 groups, respectively. Such heterogeneity in virus dynamics could be a confounding variable if it is associated with treatment allocation in compassionate use programs (i.e., observational studies). Subsequently, we mimicked randomized controlled trials of antivirals by simulation. An antiviral effect causing a 95% to 99% reduction in viral replication was added to the model. To be realistic, we assumed that randomization and treatment are initiated with some time lag after symptom onset. Using the duration of virus shedding as an outcome, the sample size to detect a statistically significant mean difference between the treatment and placebo groups (1:1 allocation) was 13,603 and 11,670 (when the antiviral effect was 95% and 99%, respectively) per group if all patients are enrolled regardless of timing of randomization. The sample size was reduced to 584 and 458 (when the antiviral effect was 95% and 99%, respectively) if only patients who are treated within 1 day of symptom onset are enrolled. We confirmed the sample size was similarly reduced when using cumulative viral load in log scale as an outcome. We used a conventional virus dynamics model, which may not fully reflect the detailed mechanisms of viral dynamics of SARS-CoV-2. The model needs to be calibrated in terms of both parameter settings and model structure, which would yield more reliable sample size calculation. In this study, we found that estimated association in observational studies can be biased due to large heterogeneity in viral dynamics among infected individuals, and statistically significant effect in randomized controlled trials may be difficult to be detected due to small sample size. The sample size can be dramatically reduced by recruiting patients immediately after developing symptoms. We believe this is the first study investigated the study design of clinical trials for antiviral treatment using the viral dynamics model.

60 APPLIED LIFE SCIENCES↗

Characterization of fault recovery through fault injection on FTMP

The development of fault-injection procedures and statistical analysis techniques to characterize the fault recovery of fault-tolerant systems is described. Pin-level fault-injection was conducted on a fault-tolerant microprocessor computer in order to generate data to assess the utility of current fault-injection sampling methods. The validity of common reliability-modeling assumptions concerning the statistical distribution of recovery times is investigated. A multiple comparison analysis for detecting behavior variations, and a distribution fitting for determining the best fit for the data were conducted. It is observed that the detection behavior is not homogeneous across all data sets, and that none of the factors under experimental control can account for the observed groupings of behavior. It is determined that no single distribution fits all the data sets, and that stratified random sampling and statistically robust parameter-estimation techniques are required to characterize fault detection time.

Finelli, George B.↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Forest Biomass Mapping From Lidar and Radar Synergies

The use of lidar and radar instruments to measure forest structure attributes such as height and biomass at global scales is being considered for a future Earth Observation satellite mission, DESDynI (Deformation, Ecosystem Structure, and Dynamics of Ice). Large footprint lidar makes a direct measurement of the heights of scatterers in the illuminated footprint and can yield accurate information about the vertical profile of the canopy within lidar footprint samples. Synthetic Aperture Radar (SAR) is known to sense the canopy volume, especially at longer wavelengths and provides image data. Methods for biomass mapping by a combination of lidar sampling and radar mapping need to be developed. In this study, several issues in this respect were investigated using aircraft borne lidar and SAR data in Howland, Maine, USA. The stepwise regression selected the height indices rh50 and rh75 of the Laser Vegetation Imaging Sensor (LVIS) data for predicting field measured biomass with a R(exp 2) of 0.71 and RMSE of 31.33 Mg/ha. The above-ground biomass map generated from this regression model was considered to represent the true biomass of the area and used as a reference map since no better biomass map exists for the area. Random samples were taken from the biomass map and the correlation between the sampled biomass and co-located SAR signature was studied. The best models were used to extend the biomass from lidar samples into all forested areas in the study area, which mimics a procedure that could be used for the future DESDYnI Mission. It was found that depending on the data types used (quad-pol or dual-pol) the SAR data can predict the lidar biomass samples with R2 of 0.63-0.71, RMSE of 32.0-28.2 Mg/ha up to biomass levels of 200-250 Mg/ha. The mean biomass of the study area calculated from the biomass maps generated by lidar- SAR synergy 63 was within 10% of the reference biomass map derived from LVIS data. The results from this study are preliminary, but do show the potential of the combined use of lidar samples and radar imagery for forest biomass mapping. Various issues regarding lidar/radar data synergies for biomass mapping are discussed in the paper.

Sun, Guoqing↗

Cross-national analysis of food security drivers: comparing results based on the Food Insecurity Experience Scale and Global Food Security Index

Abstract The second UN Sustainable Development Goal establishes food security as a priority for governments, multilateral organizations, and NGOs. These institutions track national-level food security performance with an array of metrics and weigh intervention options considering the leverage of many possible drivers. We studied the relationships between several candidate drivers and two response variables based on prominent measures of national food security: the 2019 Global Food Security Index (GFSI) and the Food Insecurity Experience Scale’s (FIES) estimate of the percentage of a nation’s population experiencing food security or mild food insecurity (FI ). We compared the contributions of explanatory variables in regressions predicting both response variables, and we further tested the stability of our results to changes in explanatory variable selection and in the countries included in regression model training and testing. At the cross-national level, the quantity and quality of a nation’s agricultural land were not predictive of either food security metric. We found mixed evidence that per-capita cereal production, per-hectare cereal yield, an aggregate governance metric, logistics performance, and extent of paid employment work were predictive of national food security. Household spending as measured by per-capita final consumption expenditure (HFCE) was consistently the strongest driver among those studied, alone explaining a median of 92% and 70% of variation (based on out-of-sample R 2 ) in GFSI and FI , respectively. The relative strength of HFCE as a predictor was observed for both response variables and was independent of the countries used for model training, the transformations applied to the explanatory variables prior to model training, and the variable selection technique used to specify multivariate regressions. The results of this cross-national analysis reinforce previous research supportive of a causal mechanism where, in the absence of exceptional local factors, an increase in income drives increase in food security. However, the strength of this effect varies depending on the countries included in regression model fitting. We demonstrate that using multiple response metrics, repeated random sampling of input data, and iterative variable selection facilitates a convergence of evidence approach to analyzing food security drivers.

42 ENGINEERING↗

Megadrought: A Series of Unfortunate La Niña Events?

Megadroughts are multidecadal periods of aridity more persistent than most droughts during the instrumental period. Paleoclimate evidence suggests that megadroughts occur in many parts of the world, including North America, Central America, western Europe, eastern Asia, and northern Africa. It remains unclear to what extent such megadroughts require external forcing or whether they can arise from internal climate variability alone. A novel statistical–dynamical approach is used to evaluate the possibility that such events arise solely as a function of interannual tropical sea surface temperature (SST) variations. A statistical emulator of tropical SST variations is constructed by using an empirical moving-blocks bootstrap approach that randomly samples multiyear sequences of the observational SST record. This approach preserves the power spectrum, seasonal cycle, and spatial pattern of El Niño-Southern Oscillation (ENSO) but removes longer timescale fluctuations embedded in the observational record. These resampled SST anomalies are then used to force an atmospheric model (the Community Atmosphere Model Version 5). As megadroughts emerge in this run, they should, therefore, be solely a consequence of La Niña sequences combined with internal atmospheric variability and persistence driven by soil moisture storage and other land-surface processes. We indeed find that megadroughts in this simulation have an amplitude-duration rate that is generally indistinguishable from the rate documented in paleoclimate records of the western United States. Our findings support the idea that megadroughts may occur randomly when the unforced climate system evolves freely over a sufficiently long period of time, implying that an unforced unusual but statistically plausible series of La Niña events may be sufficient to generate megadrought.

54 ENVIRONMENTAL SCIENCES↗

Potential of VIIRS Time Series Data for Aiding the USDA Forest Service Early Warning System for Forest Health Threats: A Gypsy Moth Defoliation Case Study

This report details one of three experiments performed during FY 2007 for the NASA RPC (Rapid Prototyping Capability) at Stennis Space Center. This RPC experiment assesses the potential of VIIRS (Visible/Infrared Imager/Radiometer Suite) and MODIS (Moderate Resolution Imaging Spectroradiometer) data for detecting and monitoring forest defoliation from the non-native Eurasian gypsy moth (Lymantria dispar). The intent of the RPC experiment was to assess the degree to which VIIRS data can provide forest disturbance monitoring information as an input to a forest threat EWS (Early Warning System) as compared to the level of information that can be obtained from MODIS data. The USDA Forest Service (USFS) plans to use MODIS products for generating broad-scaled, regional monitoring products as input to an EWS for forest health threat assessment. NASA SSC is helping the USFS to evaluate and integrate currently available satellite remote sensing technologies and data products for the EWS, including the use of MODIS products for regional monitoring of forest disturbance. Gypsy moth defoliation of the mid-Appalachian highland region was selected as a case study. Gypsy moth is one of eight major forest insect threats listed in the Healthy Forest Restoration Act (HFRA) of 2003; the gypsy moth threatens eastern U.S. hardwood forests, which are also a concern highlighted in the HFRA of 2003. This region was selected for the project because extensive gypsy moth defoliation occurred there over multiple years during the MODIS operational period. This RPC experiment is relevant to several nationally important mapping applications, including agricultural efficiency, coastal management, ecological forecasting, disaster management, and carbon management. In this experiment, MODIS data and VIIRS data simulated from MODIS were assessed for their ability to contribute broad, regional geospatial information on gypsy moth defoliation. Landsat and ASTER (Advanced Spaceborne Thermal Emission and Reflection Radiometer) data were used to assess the quality of gypsy moth defoliation mapping products derived from MODIS data and from simulated VIIRS data. The project focused on use of data from MODIS Terra as opposed to MODIS Aqua mainly because only MODIS Terra data was collected during 2000 and 2001-years with comparatively high amounts of gypsy moth defoliation within the study area. The project assessed the quality of VIIRS data simulation products. Hyperion data was employed to assess the quality of MODIS-based VIIRS simulation datasets using image correlation analysis techniques. The ART (Application Research Toolbox) software was used for data simulation. Correlation analysis between MODIS-simulated VIIRS data and Hyperion-simulated VIIRS data for red, NIR (near-infrared), and NDVI (Normalized Difference Vegetation Index) image data products collectively indicate that useful, effective VIIRS simulations can be produced using Hyperion and MODIS data sources. The r(exp 2) for red, NIR, and NDVI products were 0.56, 0.63, and 0.62, respectively, indicating a moderately high correlation between the 2 data sources. Temporal decorrelation from different data acquisition times and image misregistration may have lowered correlation results. The RPC experiment also generated MODIS-based time series data products using the TSPT (Time Series Product Tool) software. Time series of simulated VIIRS NDVI products were produced at approximately 400-meter resolution GSD (Ground Sampling Distance) at nadir for comparison to MODIS NDVI products at either 250- or 500-meter GSD. The project also computed MODIS (MOD02) NDMI (Normalized Difference Moisture Index) products at 500-meter GSD for comparison to NDVI-based products. For each year during 2000-2006, MODIS and VIIRS (simulated from MOD02) time series were computed during the peak gypsy moth defoliation time frame in the study area (approximately June 10 through July 27). Gypsy moth defoliation mapping products from simated VIIRS and MOD02 time series were produced using multiple methods, including image classification and change detection via image differencing. The latter enabled an automated defoliation detection product computed using percent change in maximum NDVI for a peak defoliation period during 2001 compared to maximum NDVI across the entire 2000-2006 time frame. Final gypsy moth defoliation mapping products were assessed for accuracy using randomly sampled locations found on available geospatial reference data (Landsat and ASTER data in conjunction with defoliation map data from the USFS). Extensive gypsy moth defoliation patches were evident on screen displays of multitemporal color composites derived from MODIS data and from simulated VIIRS vegetation index data. Such defoliation was particularly evident for 2001, although widespread denuded forests were also seen for 2000 and 2003. These visualizations were validated using aforementioned reference data. Defoliation patches were visible on displays of MODIS-based NDVI and NDMI data. The viewing of apparent defoliation patches on all of these products necessitated adoption of a specialized temporal data processing method (e.g., maximum NDVI during the peak defoliation time frame). The frequency of cloud cover necessitated this approach. Multitemporal simulated VIIRS and MODIS Terra data both produced effective general classifications of defoliated forest versus other land cover. For 2001, the MOD02-simulated VIIRS 400-meter NDVI classification produced a similar yet slightly lower overall accuracy (87.28 percent with 0.72 Kappa) than the MOD02 250-meter NDVI classification (88.44 percent with 0.75 Kappa). The MOD13 250-meter NDVI classification had a lower overall accuracy (79.13 percent) and a much lower Kappa (0.46). The report discusses accuracy assessment results in much more detail, comparing overall classification and individual class accuracy statistics for simulated VIIRS 400-meter NDVI, MOD02 250-meter NDVI, MOD02-500 meter NDVI, MOD13 250-meter NDVI, and MOD02 500-meter NDMI classifications. Automated defoliation detection products from simulated VIIRS and MOD02 data for 2001 also yielded similar, relatively high overall classification accuracy (85.55 percent for the VIIRS 400-meter NDVI versus 87.28 percent for the MOD02 250-meter NDVI). In contrast, the USFS aerial sketch map of gypsy moth defoliation showed a lower overall classification accuracy at 73.64 percent. The overall classification Kappa values were also similar for the VIIRS (approximately 0.67 Kappa) versus the MOD02 (approximately 0.72 Kappa) automated defoliation detection product, which were much higher than the values exhibited by the USFS sketch map product (overall Kappa of approximately 0.47). The report provides additional details on the accuracy of automated gypsy moth defoliation detection products compared with USFS sketch maps. The results suggest that VIIRS data can be effectively simulated from MODIS data and that VIIRS data will produce gypsy moth defoliation mapping products that are similar to MODIS-based products. The results of the RPC experiment indicate that VIIRS and MODIS data products have good potential for integration into the forest threat EWS. The accuracy assessment was performed only for 2001 because of time constraints and a relative scarcity of cloud-free Landsat and ASTER data for the peak defoliation period of the other years in the 2000-2006 time series. Additional work should be performed to assess the accuracy of gypsy moth defoliation detection products for additional years.The study area (mid-Appalachian highlands) and application (gypsy moth forest defoliation) are not necessarily representative of all forested regions and of all forest threat disturbance agents. Additional work should be performed on other inland and coastal regions as well as for other major forest threats.

Spruce, Joseph P.↗

Statistics of base polytopes in F-theory

We propose a new statistical ensemble of toric bases for elliptic Calabi-Yaus used in F-theory models, by focusing on only the convex hull of the base, i.e., the base polytope. This physically motivated coarse-graining greatly simplifies the combinatorial complexity of the part of the 4d F-theory landscape with toric bases. We develop a Monte Carlo approach that randomly samples the base polytopes within fixed boxes, with proper statistical weights. We first apply the algorithm to the set of 2d base polytopes, generating an enlarged set of toric 2d bases that include certain types of codimension-two (4,6) points, and we validate our approach against exact numbers. We then explore the set of 3d base polytopes which fit in a set of “maximal” 3d boxes, and estimate the total number of inequivalent 3d base polytopes to be 10 85 –10 90 . We provide statistical data such as the distribution of non-Higgsable gauge groups on these bases. Amusingly, a similar method can also be applied to generate reflexive polytopes in various dimensions. In both the reflexive and base polytope cases, the number of relevant polytopes obeys a Gaussian distribution as a function of the number of vertices, which can be understood in terms of other results on random polytopes in the math literature.

Differential and algebraic geometry↗

Real classical shadows

Efficiently learning expectation values of a quantum state using classical shadow tomography has become a fundamental task in quantum information theory. In a classical shadows protocol, one measures a state in a chosen basis $\mathcal{W}$ after it has evolved under a unitary transformation randomly sampled from a chosen distribution $\mathcal{U}$. In this work we study the case where $\mathcal{U}$ corresponds to either local or global orthogonal Clifford gates, and $\mathcal{W}$ consists of real-valued vectors. Our results show that for various situations of interest, this ‘real’ classical shadow protocol improves the sample complexity over the standard scheme based on general Clifford unitaries. For example, when one is interested in estimating the expectation values of arbitrary real-valued observables, global orthogonal Cliffords typically decrease the required number of samples by a factor of two. More dramatically, for k-local observables composed only of real-valued Pauli operators, sampling local orthogonal Cliffords leads to a reduction by an exponential-in-k factor in the sample complexity over local unitary Cliffords. Finally, we show that by measuring in a basis containing complex-valued vectors, orthogonal shadows can, in the limit of large system size, exactly reproduce the original unitary shadows protocol.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Similarity Downselection: Finding the n Most Dissimilar Molecular Conformers for Reference-Free Metabolomics

Computational methods for creating in silico libraries of molecular descriptors (e.g., collision cross sections) are becoming increasingly prevalent due to the limited number of authentic reference materials available for traditional library building. These so-called “reference-free metabolomics” methods require sampling sets of molecular conformers in order to produce high accuracy property predictions. Due to the computational cost of the subsequent calculations for each conformer, there is a need to sample the most relevant subset and avoid repeating calculations on conformers that are nearly identical. The goal of this study is to introduce a heuristic method of finding the most dissimilar conformers from a larger population in order to help speed up reference-free calculation methods and maintain a high property prediction accuracy. Finding the set of the n items most dissimilar from each other out of a larger population becomes increasingly difficult and computationally expensive as either n or the population size grows large. Because there exists a pairwise relationship between each item and all other items in the population, finding the set of the n most dissimilar items is different than simply sorting an array of numbers. For instance, if you have a set of the most dissimilar n = 4 items, one or more of the items from n = 4 might not be in the set n = 5. An exact solution would have to search all possible combinations of size n in the population exhaustively. We present an open-source software called similarity downselection (SDS), written in Python and freely available on GitHub. SDS implements a heuristic algorithm for quickly finding the approximate set(s) of the n most dissimilar items. We benchmark SDS against a Monte Carlo method, which attempts to find the exact solution through repeated random sampling. We show that for SDS to find the set of n most dissimilar conformers, our method is not only orders of magnitude faster, but it is also more accurate than running Monte Carlo for 1,000,000 iterations, each searching for set sizes n = 3–7 out of a population of 50,000. We also benchmark SDS against the exact solution for example small populations, showing that SDS produces a solution close to the exact solution in these instances. Using theoretical approaches, we also demonstrate the constraints of the greedy algorithm and its efficacy as a ratio to the exact solution.

97 MATHEMATICS AND COMPUTING↗

Modeled and observed properties related to the direct aerosol radiative effect of biomass burning aerosol over the southeastern Atlantic

Biomass burning smoke is advected over the southeastern Atlantic Ocean between July and October of each year. This smoke plume overlies and mixes into a region of persistent low marine clouds. Model calculations of climate forcing by this plume vary significantly in both magnitude and sign. NASA EVS-2 (Earth Venture Suborbital-2) ORACLES (ObseRvations of Aerosols above CLouds and their intEractionS) had deployments for field campaigns off the west coast of Africa in 3 consecutive years (September 2016, August 2017, and October 2018) with the goal of better characterizing this plume as a function of the monthly evolution by measuring the parameters necessary to calculate the direct aerosol radiative effect. Here, this dataset and satellite retrievals of cloud properties are used to test the representation of the smoke plume and the underlying cloud layer in two regional models (WRF-CAM5 and CNRM-ALADIN) and two global models (GEOS and UM-UKCA). The focus is on the comparisons of those aerosol and cloud properties that are the primary determinants of the direct aerosol radiative effect and on the vertical distribution of the plume and its properties. The representativeness of the observations to monthly averages are tested for each field campaign, with the sampled mean aerosol light extinction generally found to be within 20 % of the monthly mean at plume altitudes. When compared to the observations, in all models, the simulated plume is too vertically diffuse and has smaller vertical gradients, and in two of the models (GEOS and UM-UKCA), the plume core is displaced lower than in the observations. Plume carbon monoxide, black carbon, and organic aerosol masses indicate underestimates in modeled plume concentrations, leading, in general, to underestimates in mid-visible aerosol extinction and optical depth. Biases in mid-visible single scatter albedo are both positive and negative across the models. Observed vertical gradients in single scatter albedo are not captured by the models, but the models do capture the coarse temporal evolution, correctly simulating higher values in October (2018) than in August (2017) and September (2016). Uncertainties in the measured absorption Ångstrom exponent were large but propagate into a negligible (<4 %) uncertainty in integrated solar absorption by the aerosol and, therefore, in the aerosol direct radiative effect. Model biases in cloud fraction, and, therefore, the scene albedo below the plume, vary significantly across the four models. The optical thickness of clouds is, on average, well simulated in the WRF-CAM5 and ALADIN models in the stratocumulus region and is underestimated in the GEOS model; UM-UKCA simulates cloud optical thickness that is significantly too high. Overall, the study demonstrates the utility of repeated, semi-random sampling across multiple years that can give insights into model biases and how these biases affect modeled climate forcing. The combined impact of these aerosol and cloud biases on the direct aerosol radiative effect (DARE) is estimated using a first-order approximation for a subset of five comparison grid boxes. A significant finding is that the observed grid box average aerosol and cloud properties yield a positive (warming) aerosol direct radiative effect for all five grid boxes, whereas DARE using the grid-box-averaged modeled properties ranges from much larger positive values to small, negative values. It is shown quantitatively how model biases can offset each other, so that model improvements that reduce biases in only one property (e.g., single scatter albedo but not cloud fraction) would lead to even greater biases in DARE. Across the models, biases in aerosol extinction and in cloud fraction and optical depth contribute the largest biases in DARE, with aerosol single scatter albedo also making a significant contribution.

54 ENVIRONMENTAL SCIENCES↗

A software‐defined networks ‐based measurement method of network traffic for 6G technologies

Abstract In software‐defined networks (SDN) for 6G technologies, the controller should accurately and efficiently measure the flow traffic in switches for traffic engineering. Fine‐grained flow measurement can more accurately describe the traffic in the network, but it also consumes much more resources. To reduce the overhead incurred in the measurement process and obtain the approximate fine‐grained measurements, we propose a novel lightweight measurement scheme that runs in the controller. The novel lightweight measurement architecture consists of two parts: coarse‐grained measurement and interpolation‐optimization. In the first part, based on the SDN architecture, we use the pull‐based random sampling method to quickly obtain the coarse‐grained measurement of flow traffic through OpenFlow protocol. In the second part, we insert some discrete values into the coarse‐grained measurement with the interpolation theory, then we optimize interpolation results until finding the optimal fine‐grained flow traffic measurement by utilizing the multiconstraint method. We verified the feasibility of the proposed measurement method, and simulation results show that the measurement error of the proposed method is under 25%, but the flow measurement overhead of this method occupies only 3.3% compared with that of the fine‐grained method.

Huo, Liuwei↗

HopBox: An image analysis pipeline to characterize hop cone morphology

Abstract Hop cone morphology can influence picking and drying ability, and color can impact consumer preference and may be indicative of quality. However, these characteristics are not generally evaluated in hop breeding programs due to the tedious nature of trait quantification and the extensive variation among cones within a genotype. We developed the HopBox, which is a simply constructed light box with a camera mount, and a publicly available image processing pipeline that identifies hop cones within color‐corrected images, reads a QR code within the image, and outputs data on hop cone length, width, area, perimeter, openness, weight, color, and density. The trained model was applied to images of 500 cones each from 15 replicated advanced hop genotypes from the USDA‐ARS breeding program in Prosser, Washington. Analysis of variance revealed significant ( p < 0.001) differences between genotypes for all traits measured, enabling breeders to discriminate between genotypes for selection purposes. Broad sense heritability for all traits ranged from 0.23 to 0.59. A random sampling of hop cones from the complete dataset revealed that imaging only 5–10 cones adequately captured genotypic variation and provided acceptable rank correlations ( r s > 0.75); however, increasing the sample size to 30 provided optimal precision. Instructions for constructing a HopBox and the code for the analysis pipeline are publicly available online and have wide applicability for hop breeding and research.

Altendorf, Kayla R.↗

Dimension-adaptive machine learning-based quantum state reconstruction

Here, we introduce an approach for performing quantum state reconstruction on systems of n qubits using a machine learning-based reconstruction system trained exclusively on m qubits, where m ≥ n. This approach removes the necessity of exactly matching the dimensionality of a system under consideration with the dimension of a model used for training. We demonstrate our technique by performing quantum state reconstruction on randomly sampled systems of one, two, and three qubits using machine learning-based methods trained exclusively on systems containing at least one additional qubit. The reconstruction time required for machine learning-based methods scales significantly more favorably than the training time; hence this technique can offer an overall saving of resources by leveraging a single neural network for dimension-variable state reconstruction, obviating the need to train dedicated machine learning systems for each Hilbert space.

42 ENGINEERING↗

Active learning for the design of polycrystalline textures using conditional normalizing flows

Generative modeling has opened new avenues for solving previously intractable materials design problems. However, these new opportunities are accompanied by a drastic increase in the required amount of training data. This is in stark juxtaposition to the high expense and difficulty in curating such large materials datasets. In this work, we propose a novel framework for integrating generative models within an active learning loop. Further, this enables the training of generative models with datasets significantly smaller than what has previously been demonstrated, providing a direct route for their application in data constrained environments. The functionality of this framework is then demonstrated by addressing the challenge of designing polycrystalline textures associated with target anisotropic mechanical properties. The developed protocol exhibited a cost reduction between 14 to 18 times over a randomly sampled experimental design.

36 MATERIALS SCIENCE↗

Uncertainty quantification of a physics-informed model based on sparse identification of a Thermal Energy Distribution System

Integrated energy systems (IES)s are crucial for enhancing the economy and efficiency of power generation sources (e.g., nuclear energy) necessary to unleash American energy dominance. These systems can be integrated with thermal energy storage (TES) and intermittent renewable energies to optimize overall energy use, peak-load regulation, and demand-side responses. However, the stabilization of energy generation, transport, and utilization introduces operational complexities that exceed the challenges of managing each sub-component individually. Currently, though IESs rely on human operators for efficiency and stability, reducing human error risk and enhancing performance through automation is highly desirable. Recent advances at Idaho National Laboratory have demonstrated successful control of the Thermal Energy Distributed System (TEDS). However, the automatic control system depends on a deterministic Sparse Identification of Nonlinear Dynamics with Control (SINDyC) model, which are trained based on simulation data from physics-based simulations. Because of uncertainties in physics-based simulation, SINDyC model results in large discrepancies against experimental data and cannot be reliably used in automatic control. In this paper, we present an innovative approach to address these discrepancies by quantifying uncertainties and developing a more robust model. We first generated trajectories by using first-principles physics codes to encapsulate the experiment. Next, we trained thousands of models by randomly sampling these trajectories. We then collapsed all those models into one probabilistic SINDyC by fitting a multivariate Gaussian distribution onto the resulting coefficient’s distribution. Despite its simplicity, our approach successfully produced 95% confidence intervals that captured the experimental trajectories. It even did so with a higher probability and better U-pooling score across six of the seven relevant quantities of interest (QoIs), as compared to other classical approaches. In conclusion, ongoing research is focusing on generating new experimental trajectories to validate this approach, and on employing Bayesian calibration to refine parametric uncertainties and guide future model development efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Piston geometry and stroke optimization for high efficiency propane spark ignition engines

Propane has unique properties and offers interesting characteristics for high-efficiency spark ignition engines. Its high volatility reduces or completely eliminates fuel-wall wetting and facilitates fuel air mixing. Furthermore, propane has a research octane number of 112 and a high octane sensitivity of 15. Finally, its laminar flame speed is on the same order as that of conventional gasoline, and it exhibits high dilution tolerance. Modern spark ignition internal combustion engines rely on fast combustion rates and high dilution to achieve high brake thermal efficiencies. To accomplish this, high stroke-to-bore ratios and high geometric compression ratios have been used in new engine designs. Therefore, propane’s relatively high laminar flame speeds, high knock resistance, and dilution tolerance make it an excellent candidate fuel for modern spark ignition engines. The objective of this work is to co-optimize the piston geometry and the engine stroke to maximize the efficiency of a spark-ignition engine fueled with propane. 3D computational fluid dynamics (CFD) simulations employing the extended coherent flamelet model were used to study the parametric effects of piston shape and stroke length. A piston geometry based on high performing pistons was parameterized using four controlling parameters. The piston geometry and engine stroke design space was explored using deterministic and quasi-random sampling techniques. In conclusion, a Gaussian process regression model was built using the simulation data to explain the results observed.

33 ADVANCED PROPULSION SYSTEMS↗