Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data imbalance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Geography program, design, structure and operational strategy

The geography program is designed to move systematically toward a capability to increase remote sensing data into operational systems for monitoring land use and related environmental change. The problems of environmental imbalance arising from rapid urbanization and other dramatic changes in land use are considered. These overall problems translate into working level problems of establishing the validity of various sensor-data combinations that will best obtain the regional land use and environmental information. The goal, to better understand, predict, and assist policy makers to regulate urban and regional land use changes resulting from population growth and technological advancement, is put forth.

Alexander, R. H.↗

Feature Acquisition with Imbalanced Training Data

This work considers cost-sensitive feature acquisition that attempts to classify a candidate datapoint from incomplete information. In this task, an agent acquires features of the datapoint using one or more costly diagnostic tests, and eventually ascribes a classification label. A cost function describes both the penalties for feature acquisition, as well as misclassification errors. A common solution is a Cost Sensitive Decision Tree (CSDT), a branching sequence of tests with features acquired at interior decision points and class assignment at the leaves. CSDT's can incorporate a wide range of diagnostic tests and can reflect arbitrary cost structures. They are particularly useful for online applications due to their low computational overhead. In this innovation, CSDT's are applied to cost-sensitive feature acquisition where the goal is to recognize very rare or unique phenomena in real time. Example applications from this domain include four areas. In stream processing, one seeks unique events in a real time data stream that is too large to store. In fault protection, a system must adapt quickly to react to anticipated errors by triggering repair activities or follow- up diagnostics. With real-time sensor networks, one seeks to classify unique, new events as they occur. With observational sciences, a new generation of instrumentation seeks unique events through online analysis of large observational datasets. This work presents a solution based on transfer learning principles that permits principled CSDT learning while exploiting any prior knowledge of the designer to correct both between-class and withinclass imbalance. Training examples are adaptively reweighted based on a decomposition of the data attributes. The result is a new, nonparametric representation that matches the anticipated attribute distribution for the target events.

Thompson, David R.↗

Dynamic Load Balancing For Grid Partitioning on a SP-2 Multiprocessor: A Framework

Computational requirements of full scale computational fluid dynamics change as computation progresses on a parallel machine. The change in computational intensity causes workload imbalance of processors, which in turn requires a large amount of data movement at runtime. If parallel CFD is to be successful on a parallel or massively parallel machine, balancing of the runtime load is indispensable. Here a framework is presented for dynamic load balancing for CFD applications, called Jove. One processor is designated as a decision maker Jove while others are assigned to computational fluid dynamics. Processors running CFD send flags to Jove in a predetermined number of iterations to initiate load balancing. Jove starts working on load balancing while other processors continue working with the current data and load distribution. Jove goes through several steps to decide if the new data should be taken, including preliminary evaluate, partition, processor reassignment, cost evaluation, and decision. Jove running on a single IBM SP2 node has been completely implemented. Preliminary experimental results show that the Jove approach to dynamic load balancing can be effective for full scale grid partitioning on the target machine IBM SP2.

Sohn, Andrew↗

Dynamic Load Balancing for Finite Element Calculations on Parallel Computers

Computational requirements of full scale computational fluid dynamics change as computation progresses on a parallel machine. The change in computational intensity causes workload imbalance of processors, which in turn requires a large amount of data movement at runtime. If parallel CFD is to be successful on a parallel or massively parallel machine, balancing of the runtime load is indispensable. Here a frame work is presented for dynamic load balancing for CFD applications, called Jove. One processor is designated as a decision maker Jove while others are assigned to computational fluid dynamics. Processors running CFD send flags to Jove in a predetermined number of iterations to initiate load balancing. Jove starts working on load balancing while other processors continue working with the current data and load distribution. Jove goes through several steps to decide if the new data should be taken, including preliminary evaluate, partition, processor reassignment, cost evaluation, and decision. Jove running on a single SP2 node has been completely implemented. Preliminary experimental results show that the Jove approach to dynamic load balancing can be effective for full scale grid partitioning on the target machine SP2.

Pramono, Eddy↗

Dynamic Load Balancing for Grid Partitioning on a SP-2 Multiprocessor: A Framework

Computational requirements of full scale computational fluid dynamics change as computation progresses on a parallel machine. The change in computational intensity causes workload imbalance of processors, which in turn requires a large amount of data movement at runtime. If parallel CFD is to be successful on a parallel or massively parallel machine, balancing of the runtime load is indispensable. Here a framework is presented for dynamic load balancing for CFD applications, called Jove. One processor is designated as a decision maker Jove while others are assigned to computational fluid dynamics. Processors running CFD send flags to Jove in a predetermined number of iterations to initiate load balancing. Jove starts working on load balancing while other processors continue working with the current data and load distribution. Jove goes through several steps to decide if the new data should be taken, including preliminary evaluate, partition, processor reassignment, cost evaluation, and decision. Jove running on a single EBM SP2 node has been completely implemented. Preliminary experimental results show that the Jove approach to dynamic load balancing can be effective for full scale grid partitioning on the target machine IBM SP2.

Sohn, Andrew↗

Data Mining for Understanding and Improving Decision-making Affecting Ground Delay Programs

The continuous growth in the demand for air transportation results in an imbalance between airspace capacity and traffic demand. The airspace capacity of a region depends on the ability of the system to maintain safe separation between aircraft in the region. In addition to growing demand, the airspace capacity is severely limited by convective weather. During such conditions, traffic managers at the FAA's Air Traffic Control System Command Center (ATCSCC) and dispatchers at various Airlines' Operations Center (AOC) collaborate to mitigate the demand-capacity imbalance caused by weather. The end result is the implementation of a set of Traffic Flow Management (TFM) initiatives such as ground delay programs, reroute advisories, flow metering, and ground stops. Data Mining is the automated process of analyzing large sets of data and then extracting patterns in the data. Data mining tools are capable of predicting behaviors and future trends, allowing an organization to benefit from past experience in making knowledge-driven decisions.

Weather↗

Data Mining for Understanding and Impriving Decision-Making Affecting Ground Delay Programs

The continuous growth in the demand for air transportation results in an imbalance between airspace capacity and traffic demand. The airspace capacity of a region depends on the ability of the system to maintain safe separation between aircraft in the region. In addition to growing demand, the airspace capacity is severely limited by convective weather. During such conditions, traffic managers at the FAA's Air Traffic Control System Command Center (ATCSCC) and dispatchers at various Airlines' Operations Center (AOC) collaborate to mitigate the demand-capacity imbalance caused by weather. The end result is the implementation of a set of Traffic Flow Management (TFM) initiatives such as ground delay programs, reroute advisories, flow metering, and ground stops. Data Mining is the automated process of analyzing large sets of data and then extracting patterns in the data. Data mining tools are capable of predicting behaviors and future trends, allowing an organization to benefit from past experience in making knowledge-driven decisions. The work reported in this paper is focused on ground delay programs. Data mining algorithms have the potential to develop associations between weather patterns and the corresponding ground delay program responses. If successful, they can be used to improve and standardize TFM decision resulting in better predictability of traffic flows on days with reliable weather forecasts. The approach here seeks to develop a set of data mining and machine learning models and apply them to historical archives of weather observations and forecasts and TFM initiatives to determine the extent to which the theory can predict and explain the observed traffic flow behaviors.

data mining↗

Global Carbon Budget 2022

Accurate assessment of anthropogenic carbon dioxide (CO 2 ) emissions and their redistribution among the atmosphere, ocean, and terrestrial biosphere in a changing climate is critical to better understand the global carbon cycle, support the development of climate policies, and project future climate change. Here we describe and synthesize data sets and methodologies to quantify the five major components of the global carbon budget and their uncertainties. Fossil CO 2 emissions (E FOS ) are based on energy statistics and cement production data, while emissions from land-use change (E LUC ), mainly deforestation, are based on land use and land-use change data and bookkeeping models. Atmospheric CO 2 concentration is measured directly, and its growth rate (G ATM ) is computed from the annual changes in concentration. The ocean CO 2 sink (S OCEAN ) is estimated with global ocean biogeochemistry models and observation-based data products. The terrestrial CO 2 sink (S LAND ) is estimated with dynamic global vegetation models. The resulting carbon budget imbalance (B IM ), the difference between the estimated total emissions and the estimated changes in the atmosphere, ocean, and terrestrial biosphere, is a measure of imperfect data and understanding of the contemporary carbon cycle. All uncertainties are reported as ±1σ. For the year 2021, E FOS increased by 5.1 % relative to 2020, with fossil emissions at 10.1 ± 0.5 GtC yr −1 (9.9 ± 0.5 GtC yr −1 when the cement carbonation sink is included), and E LUC was 1.1 ± 0.7 GtC yr −1 , for a total anthropogenic CO 2 emission (including the cement carbonation sink) of 10.9 ± 0.8 GtC yr −1 (40.0 ± 2.9 GtCO 2 ). Also, for 2021, G ATM was 5.2 ± 0.2 GtC yr −1 (2.5 ± 0.1 ppm yr −1 ), S OCEAN was 2.9 ± 0.4 GtC yr −1 , and S LAND was 3.5 ± 0.9 GtC yr −1 , with a B IM of −0.6 GtC yr −1 (i.e. the total estimated sources were too low or sinks were too high). The global atmospheric CO 2 concentration averaged over 2021 reached 414.71 ± 0.1 ppm. Preliminary data for 2022 suggest an increase in E FOS relative to 2021 of +1.0 % (0.1 % to 1.9 %) globally and atmospheric CO 2 concentration reaching 417.2 ppm, more than 50 % above pre-industrial levels (around 278 ppm). Overall, the mean and trend in the components of the global carbon budget are consistently estimated over the period 1959–2021, but discrepancies of up to 1 GtC yr −1 persist for the representation of annual to semi-decadal variability in CO 2 fluxes. Comparison of estimates from multiple approaches and observations shows (1) a persistent large uncertainty in the estimate of land-use change emissions, (2) a low agreement between the different methods on the magnitude of the land CO 2 flux in the northern extratropics, and (3) a discrepancy between the different methods on the strength of the ocean sink over the last decade. This living data update documents changes in the methods and data sets used in this new global carbon budget and the progress in understanding of the global carbon cycle compared with previous publications of this data set. The data presented in this work are available at https://doi.org/10.18160/GCP-2022 (Friedlingstein et al., 2022b).

Pierre Friedlingstein↗

Low luminosity radio galaxies - Effects of gaseous environment

X-ray data from the Einstein Observatory have been collected for a number of low luminosity B2 radio galaxies, in order to study the effect of gas pressure on radio sources. Generally, the thermal pressure of the X-ray emitting gas is sufficient to confine most of the radio components. Various possibilities are given for sources where thermal pressure significantly exceeds the nonthermal pressure. For radio lobes, an inverse correlation is found between the central densities of the cluster gas and the size of the radio structures. For jets, a study of the pressure ratios is given for different distances from the radio core. In a few cases, pressure imbalance is found in the outer regions of the jets. However, from depolarization data, this is unlikely to be produced by a large thermal component inside the radio jets.

Morganti, R.↗

Tailoff thrust and impulse imbalance between pairs of Space Shuttle solid rocket motors

The tailoff thrust and impulse imbalance between pairs of solid rocket motors is of particular interest for the Space Shuttle Vehicle because of the potential control problems that exist with this asymmetric configuration. Although a similar arrangement of solid rocket motors was utilized for the Titan Program, they produced less than one-half the thrust level of the Space Shuttle at web action time, and the overall vehicle was symmetric. Since the Titan Program does provide the most applicable actual test data, 23 flight pairs were analyzed to determine the actual tailoff thrust and impulse imbalance experienced. The results were scaled up using the predicted web action time thrust and tailoff time to arrive at values for the Space Shuttle. These values were then statistically treated to obtain a prediction of the maximum imbalance one could expect to experience during the Shuttle Program.

Jacobs, E. P.↗

Monte Carlo investigation of thrust imbalance of solid rocket motor pairs

The Monte Carlo method of statistical analysis is used to investigate the theoretical thrust imbalance of pairs of solid rocket motors (SRMs) firing in parallel. Sets of the significant variables are selected using a random sampling technique and the imbalance calculated for a large number of motor pairs using a simplified, but comprehensive, model of the internal ballistics. The treatment of burning surface geometry allows for the variations in the ovality and alignment of the motor case and mandrel as well as those arising from differences in the basic size dimensions and propellant properties. The analysis is used to predict the thrust-time characteristics of 130 randomly selected pairs of Titan IIIC SRMs. A statistical comparison of the results with test data for 20 pairs shows the theory underpredicts the standard deviation in maximum thrust imbalance by 20% with variability in burning times matched within 2%. The range in thrust imbalance of Space Shuttle type SRM pairs is also estimated using applicable tolerances and variabilities and a correction factor based on the Titan IIIC analysis.

Sforzini, R. H.↗

Improved Groundwater Table and L-Band Brightness Temperature Estimates for Northern Hemisphere Peatlands Using New Model Physics and SMOS Observations in a Global Data Assimilation Framework

There is an urgent need to include northern peatland hydrology in global Earth system models to better understand land-atmosphere interactions and sensitivities of peatland functions to climate change, and, ultimately, to improve climate change predictions. In this study, we introduced for the first time peatland-specific model physics into an assimilation scheme for L-band brightness temperature (Tb) data from the Soil Moisture Ocean Salinity (SMOS) mission to improve groundwater table estimates. We conducted two sets of model-only and data assimilation experiments using the Catchment Land Surface Model (CLSM), applying (over peatlands only) in one of them a peatland-specific adaptation (PEATCLSM). The evaluation against in-situ measurements of peatland groundwater table depth indicates the superiority of PEATCLSM model physics and additionally improved performance after assimilating SMOS Tb observations. The better performance of PEATCLSM over nearly all Northern Hemisphere peatlands is further supported by the better agreement between SMOS Tb observations and Tb estimates from the model-only and data assimilation runs. Within the data assimilation scheme, PEATCLSM reduces Tb observation-minus-forecast residuals and leads to reduced data assimilation updates of water storage components and, thus, reduced water budget imbalances in the assimilation system.

SMOS↗

Efficient Load Balancing and Data Remapping for Adaptive Grid Calculations

Mesh adaption is a powerful tool for efficient unstructured- grid computations but causes load imbalance among processors on a parallel machine. We present a novel method to dynamically balance the processor workloads with a global view. This paper presents, for the first time, the implementation and integration of all major components within our dynamic load balancing strategy for adaptive grid calculations. Mesh adaption, repartitioning, processor assignment, and remapping are critical components of the framework that must be accomplished rapidly and efficiently so as not to cause a significant overhead to the numerical simulation. Previous results indicated that mesh repartitioning and data remapping are potential bottlenecks for performing large-scale scientific calculations. We resolve these issues and demonstrate that our framework remains viable on a large number of processors.

Oliker, Leonid↗

The Persistent Challenge of Data Locality in the Post-Exascale Era

The era of exascale computing, exemplified by systems like Frontier achieving exaflop-level performance, marks a milestone. However, the quest for sheer compute power leads to strong imbalance in system design. Hence, scaling advancements in memory, network bandwidth, and storage are also necessary and pose challenges, with a crucial need to address data locality issues. This article underscores the fundamental importance of data locality as a key abstraction for optimizing application performance. Despite notable software solutions, the growing complexity of parallelism and memory hierarchy demands performance-portable data locality solutions across diverse computing platforms. Additionally, the article revisits data locality aspects, covering hardware considerations, application perspectives, software stack abstractions, and tool support. It concludes with insights into data locality challenges and opportunities, emphasizing the ongoing significance of collaborative research for progress in this critical issue.

Unat, Didem [Koc University, Istanbul (Turkey)] (O↗

First Year Wilkinson Microwave Anisotropy Probe(WMAP) Observations: Data Processing Methods and Systematic Errors Limits

We describe the calibration and data processing methods used to generate full-sky maps of the cosmic microwave background (CMB) from the first year of Wilkinson Microwave Anisotropy Probe (WMAP) observations. Detailed limits on residual systematic errors are assigned based largely on analyses of the flight data supplemented, where necessary, with results from ground tests. The data are calibrated in flight using the dipole modulation of the CMB due to the observatory's motion around the Sun. This constitutes a full-beam calibration source. An iterative algorithm simultaneously fits the time-ordered data to obtain calibration parameters and pixelized sky map temperatures. The noise properties are determined by analyzing the time-ordered data with this sky signal estimate subtracted. Based on this, we apply a pre-whitening filter to the time-ordered data to remove a low level of l/f noise. We infer and correct for a small (approx. 1 %) transmission imbalance between the two sky inputs to each differential radiometer, and we subtract a small sidelobe correction from the 23 GHz (K band) map prior to further analysis. No other systematic error corrections are applied to the data. Calibration and baseline artifacts, including the response to environmental perturbations, are negligible. Systematic uncertainties are comparable to statistical uncertainties in the characterization of the beam response. Both are accounted for in the covariance matrix of the window function and are propagated to uncertainties in the final power spectrum. We characterize the combined upper limits to residual systematic uncertainties through the pixel covariance matrix.

Hinshaw, G.↗

Extraction of quantitative surface characteristics from AIRSAR data for Death Valley, California

Polarimetric Airborne Synthetic Aperture Radar (AIRSAR) data were collected for the Geologic Remote Sensing Field Experiment (GRSFE) over Death Valley, California, USA, in Sep. 1989. AIRSAR is a four-look, quad-polarization, three frequency instrument. It collects measurements at C-band (5.66 cm), L-band (23.98 cm), and P-band (68.13 cm), and has a GIFOV of 10 meters and a swath width of 12 kilometers. Because the radar measures at three wavelengths, different scales of surface roughness are measured. Also, dielectric constants can be calculated from the data. The AIRSAR data were calibrated using in-scene trihedral corner reflectors to remove cross-talk; and to calibrate the phase, amplitude, and co-channel gain imbalance. The calibration allows for the extraction of accurate values of rms surface roughness, dielectric constants, sigma(sub 0) backscatter, and polarization information. The radar data sets allow quantitative characterization of small scale surface structure of geologic units, providing information about the physical and chemical processes that control the surface morphology. Combining the quantitative information extracted from the radar data with other remotely sensed data sets allows discrimination, identification and mapping of geologic units that may be difficult to discern using conventional techniques.

Kierein-Young, K. S.↗

Some dynamical design considerations for momentum biased spacecraft

This paper discusses some of the dynamical aspects that affect the attitude behavior of a spacecraft placed in a near earth polar orbit. The parameters such as atmospheric density variation, super-rotation of upper atmosphere, pitch equilibrium angle, and trim boom length selection to reduce pitch axis drift are considered. These parameters are studied with reference to two momentum biased satellities (Magsat and Dynamics Explorer-2) and the analyses are correlated with observed flight data. It is found that the design model of certain parameters has to be improved by including other phenomena so that flight data and the analytical model predictions have good correlations. The effects of mass imbalance on the orbital and nutational modes of the spacecraft attitude are discussed.

Sellappan, R. G.↗

An Atmospheric General Circulation Model with Chemistry for the CRAY T3E: Design, Performance Optimization and Coupling to an Ocean Model

The design, implementation and performance optimization on the CRAY T3E of an atmospheric general circulation model (AGCM) which includes the transport of, and chemical reactions among, an arbitrary number of constituents is reviewed. The parallel implementation is based on a two-dimensional (longitude and latitude) data domain decomposition. Initial optimization efforts centered on minimizing the impact of substantial static and weakly-dynamic load imbalances among processors through load redistribution schemes. Recent optimization efforts have centered on single-node optimization. Strategies employed include loop unrolling, both manually and through the compiler, the use of an optimized assembler-code library for special function calls, and restructuring of parts of the code to improve data locality. Data exchanges and synchronizations involved in coupling different data-distributed models can account for a significant fraction of the running time. Therefore, the required scattering and gathering of data must be optimized. In systems such as the T3E, there is much more aggregate bandwidth in the total system than in any particular processor. This suggests a distributed design. The design and implementation of a such distributed 'Data Broker' as a means to efficiently couple the components of our climate system model is described.

Farrara, John D.↗