Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher↗

Restructuring Big Data to Improve Data Access and Performance in Analytic Services Making Research More Efficient for the Study of Extreme Weather Events and Application User Communities

By developing and enhancing various services and tools, the GES DISC provides users with the capability to access and visualize data, and to make comparisons of data from multiple sensor and models via a number of cross-discipline projects. Discovering Data via Faceted Web Interface Web interface to data products and services Search and Download mechanisms Dataset Landing Pages Accessing Data through Interoperable Services: GDS – GrADS Data Server OPeNDAP - Open-source Project for a Network Data Access Protocol WMS – OGC service GIS connector – allowing IS tools to access data easier (coming soon) HTTPS -- direct online access Downloading Data Basics: Subset and egridding Service – Parameter, Spatial, Time, Vertical, Mean averaging, format conversion, and regridding for L3/L4 gridded data Swath Data Subsetter – Parameter, spatial subset of L2 /L1 data. Visualizing Data Online: Giovanni –Visualization and Analysis L3/L4 gridded data AIRS NRT Viewer – AIRS near-real-time DQVis – L2 data quality visualization

data cube↗

An Integrated Gate Turnaround Management Concept Leveraging Big Data/Analytics for NAS Performance Improvements

The Integrated Gate Turnaround Management (IGTM) concept was developed to improve the gate turnaround performance at the airport by leveraging relevant historical data to support optimization of airport gate operations, which include: taxi to the gate, gate services, push back, taxi to the runway, and takeoff, based on available resources, constraints, and uncertainties. By analyzing events of gate operations, primary performance dependent attributes of these events were identified for the historical data analysis such that performance models can be developed based on uncertainties to support descriptive, predictive, and prescriptive functions. A system architecture was developed to examine system requirements in support of such a concept. An IGTM prototype was developed to demonstrate the concept using a distributed network and collaborative decision tools for stakeholders to meet on time pushback performance under uncertainties.

Big Data and Net-enabled ATM↗

Earth Science Data Analytics: Preparing for Extracting Knowledge from Information

Data analytics is the process of examining large amounts of data of a variety of types to uncover hidden patterns, unknown correlations and other useful information. Data analytics is a broad term that includes data analysis, as well as an understanding of the cognitive processes an analyst uses to understand problems and explore data in meaningful ways. Analytics also include data extraction, transformation, and reduction, utilizing specific tools, techniques, and methods. Turning to data science, definitions of data science sound very similar to those of data analytics (which leads to a lot of the confusion between the two). But the skills needed for both, co-analyzing large amounts of heterogeneous data, understanding and utilizing relevant tools and techniques, and subject matter expertise, although similar, serve different purposes. Data Analytics takes on a practitioners approach to applying expertise and skills to solve issues and gain subject knowledge. Data Science, is more theoretical (research in itself) in nature, providing strategic actionable insights and new innovative methodologies. Earth Science Data Analytics (ESDA) is the process of examining, preparing, reducing, and analyzing large amounts of spatial (multi-dimensional), temporal, or spectral data using a variety of data types to uncover patterns, correlations and other information, to better understand our Earth. The large variety of datasets (temporal spatial differences, data types, formats, etc.) invite the need for data analytics skills that understand the science domain, and data preparation, reduction, and analysis techniques, from a practitioners point of view. The application of these skills to ESDA is the focus of this presentation. The Earth Science Information Partners (ESIP) Federation Earth Science Data Analytics (ESDA) Cluster was created in recognition of the practical need to facilitate the co-analysis of large amounts of data and information for Earth science. Thus, from a to advance science point of view: On the continuum of ever evolving data management systems, we need to understand and develop ways that allow for the variety of data relationships to be examined, and information to be manipulated, such that knowledge can be enhanced, to facilitate science. Recognizing the importance and potential impacts of the unlimited ways to co-analyze heterogeneous datasets, now and especially in the future, one of the objectives of the ESDA cluster is to facilitate the preparation of individuals to understand and apply needed skills to Earth science data analytics. Pinpointing and communicating the needed skills and expertise is new, and not easy. Information technology is just beginning to provide the tools for advancing the analysis of heterogeneous datasets in a big way, thus, providing opportunity to discover unobvious scientific relationships, previously invisible to the science eye. And it is not easy It takes individuals, or teams of individuals, with just the right combination of skills to understand the data and develop the methods to glean knowledge out of data and information. In addition, whereas definitions of data science and big data are (more or less) available (summarized in Reference 5), Earth science data analytics is virtually ignored in the literature, (barring a few excellent sources).

data analytics↗

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and data science, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Core Size↗

High‐Resolution National‐Scale Water Modeling Is Enhanced by Multiscale Differentiable Physics‐Informed Machine Learning

Abstract The National Water Model (NWM) is a key tool for flood forecasting, planning, and water management. Key challenges facing the NWM include calibration and parameter regionalization when confronted with big data. We present two novel versions of high‐resolution (∼37 km 2 ) differentiable models (a type of hybrid model): one with implicit, unit‐hydrograph‐style routing and another with explicit Muskingum‐Cunge routing in the river network. The former predicts streamflow at basin outlets whereas the latter presents a discretized product that seamlessly covers rivers in the conterminous United States (CONUS). Both versions use neural networks to provide a multiscale parameterization and process‐based equations to provide a structural backbone, which were trained simultaneously (“end‐to‐end”) on 2,807 basins across the CONUS and evaluated on 4,997 basins. Both versions show great potential to elevate future NWM performance for extensively calibrated as well as ungauged sites: the median daily Nash‐Sutcliffe efficiency of all 4,997 basins is improved to around 0.68 from 0.48 of NWM3.0. As they resolve spatial heterogeneity, both versions greatly improved simulations in the western CONUS and also in the Prairie Pothole Region, a long‐standing modeling challenge. The Muskingum‐Cunge version further improved performance for basins >10,000 km 2 . Overall, our results show how neural‐network‐based parameterizations can improve NWM performance for providing operational flood predictions while maintaining interpretability and multivariate outputs. The modeling system supports the Basic Model Interface (BMI), which allows seamless integration with the next‐generation NWM. We also provide a CONUS‐scale hydrologic data set for further evaluation and use.

Song, Yalan [Civil and Environmental Engineering T↗

NASA's Hyperwall Revealing the Big Picture

NASA:s hyperwall is a sophisticated visualization tool used to display large datasets. The hyperwall, or video wall, is capable of displaying multiple high-definition data visualizations and/or images simultaneously across an arrangement of screens. Functioning as a key component at many NASA exhibits, the hyperwall is used to help explain phenomena, ideas, or examples of world change. The traveling version of the hyperwall is typically comprised of nine 42-50" flat-screen monitors arranged in a 3x3 array (as depicted below). However, it is not limited to monitor size or number; screen sizes can be as large as 52" and the arrangement of screens can include more than nine monitors. Generally, NASA satellite and model data are used to highlight particular themes in atmospheric, land, and ocean science. Many of the existing hyperwall stories reveal change across space and time, while others display large-scale still-images accompanied by descriptive, story-telling captions. Hyperwall content on a variety of Earth Science topics already exists and is made available to the public at: eospso.gsfc.nasa.gov/hyperwall. Keynote and PowerPoint presentations as well as Summary of Story files are available for download on each existing topic. New hyperwall content and accompanying files will continue being developed to promote scientific literacy across a diverse group of audience members. NASA invites the use of content accessible through this website but requests the user to acknowledge any and all data sources referenced in the content being used.

Sellers, Piers↗

VEDA Visualization Exploration & Data Analysis

Why? - Interdisciplinary science depends on large amount of Earth science data and computational resources - Working with these datasets is non-trivial - Big data science requires advanced distributed computing knowledge What? VEDA is an open platform that brings key Earth science datasets next to open source tools for data processing, analysis, visualization, and exploration in a managed and more accessible computing environment.

Manil Maskey↗

7th World Congress on Integrated Computational Materials Engineering (ICME 2023) (Final Technical Report)

Integrated Computational Materials Engineering (ICME) has received international attention due to its potential to shorten product development time, while lowering cost and improving design and manufacturing outcomes. ICME is an approach to designing materials solutions for specific applications that use computer modeling programs to predict the behavior of materials and integrate this information into the overall materials, processing, and manufacturing design cycle. The 7th World Congress on Integrated Computational Materials Engineering (ICME 2023) was held in Orlando, Florida from May 21–25, 2023 with the goal to convene stakeholders from across all areas of modeling and simulation, experimental specialization, and design, as well as from across academia, government, and industry, to address ICME tools and techniques and their integration, as well as to examine their application in engineering. This atmosphere facilitated rich interactions between the experimentalists, modelers, and computational and design, from academia, government, and industry, to discuss ICME tools and techniques and their application in engineering.

36 MATERIALS SCIENCE↗

Contrast reduction by the atmosphere and retrieval of nonuniform surface reflectance

A radiative transfer model is developed which gives the upward radiance at nadir for any 1-D Lambertian surface reflectance. This model is used to depict the atmospheric effect on the transmittance of contrast for any 1-D surface reflectance. Here by contrast we mean a general variation of the radiation field across the image. With the aid of this model an inversion algorithm is developed for retrieval of true surface reflectance from high resolution satellite data (e.g., Landsat). This inversion technique can be a useful tool for extraction of surface reflectance from satellite data in the case of a surface reflectance variable in one dimension only (e.g., seashore or near borders of big fields). A sensitivity study of the inversion procedure on the knowledge of atmospheric parameters and sensor calibration was performed. It is shown that this inversion technique is stable even in the presence of errors in the sensor calibration and the atmospheric parameters. The method was applied to Landsat data in two wavelengths. The results show reasonable dependence of the derived surface reflectance on the distance from the seashore.

Mekler, Y.↗

BIG MAC: A bolometer array for mid-infrared astronomy, Center Director's Discretionary Fund

The infrared array referred to as Big Mac (for Marshall Array Camera), was designed for ground based astronomical observations in the wavelength range 5 to 35 microns. It contains 20 discrete gallium-doped germanium bolometer detectors at a temperature of 1.4K. Each bolometer is irradiated by a square field mirror constituting a single pixel of the array. The mirrors are arranged contiguously in four columns and five rows, thus defining the array configuration. Big Mac utilized cold reimaging optics and an up looking dewar. The total Big Mac system also contains a telescope interface tube for mounting the dewar and a computer for data acquisition and processing. Initial astronomical observations at a major infrared observatory indicate that Big Mac performance is excellent, having achieved the design specifications and making this instrument an outstanding tool for astrophysics.

Telesco, C. M.↗

Communication Network Awareness Machine System Phase I Development: The Intelligent Party-Line Schema

As NextGen continues toward the full implementation of a Net-Centric Architecture (N-CA)it will inherently provide a continuous increase to the Three-Vs components (Volume, Velocity, and Variety) of big data . This will create an insurmountable environment for direct-action aviation personnel (DAAP)as the DAAP’s natural abilities to manage and process data into actionable information will be overmatched by the Three-Vs. Therefore, conducting operations within a N-CA requires that new tools and applications be researched and developed to aid the DAAP’s ability to understand and manage data, mitigate non-normals, create contingency plans and actions. This paper will describe a research area at NASA Langley Research Center known as the Intelligent Party-Line (IPL).

Intelligent Party-Line↗

IN13B-1660: Analytics and Visualization Pipelines for Big Data on the NASA Earth Exchange (NEX) and OpenNEX

We are developing capabilities for an integrated petabyte-scale Earth science collaborative analysis and visualization environment. The ultimate goal is to deploy this environment within the NASA Earth Exchange (NEX) and OpenNEX in order to enhance existing science data production pipelines in both high-performance computing (HPC) and cloud environments. Bridging of HPC and cloud is a fairly new concept under active research and this system significantly enhances the ability of the scientific community to accelerate analysis and visualization of Earth science data from NASA missions, model outputs and other sources. We have developed a web-based system that seamlessly interfaces with both high-performance computing (HPC) and cloud environments, providing tools that enable science teams to develop and deploy large-scale analysis, visualization and QA pipelines of both the production process and the data products, and enable sharing results with the community. Our project is developed in several stages each addressing separate challenge - workflow integration, parallel execution in either cloud or HPC environments and big-data analytics or visualization. This work benefits a number of existing and upcoming projects supported by NEX, such as the Web Enabled Landsat Data (WELD), where we are developing a new QA pipeline for the 25PB system.

visualization↗

Watershed Modeling with Remotely Sensed Big Data: MODIS Leaf Area Index Improves Hydrology and Water Quality Predictions

Traditional watershed modeling often overlooks the role of vegetation dynamics. There is also little quantitative evidence to suggest that increased physical realism of vegetation dynamics in process-based models improves hydrology and water quality predictions simultaneously. In this study, we applied a modified Soil and Water Assessment Tool (SWAT) to quantify the extent of improvements that the assimilation of remotely sensed Leaf Area Index (LAI) would convey to streamflow, soil moisture, and nitrate load simulations across a 16,860 km2 agricultural watershedin the midwestern United States. We modified the SWAT source code to automatically override the model’s built-in semiempirical LAI with spatially distributed and temporally continuous estimates from Moderate Resolution Imaging Spectroradiometer (MODIS). Compared to a “basic” traditional model with limited spatial information, our LAI assimilation model (i) significantly improved daily streamflow simulations during medium-to-low flow conditions, (ii) provided realistic spatial distributions of growing season soil moisture, and (iii) substantially reproduced the long-term observed variability of daily nitrate loads. Further analysis revealed that the overestimation or underestimation of LAI imparted a proportional cascading effect on how the model partitions hydrologic fluxes and nutrient pools. As such, assimilation of MODIS LAI data corrected the model’sLAI overestimation tendency, which led to a proportionally increased rootzone soil moisture and decreased plant nitrogen uptake. With these new findings, our study fills the existing knowledge gap regarding vegetation dynamics in watershed modeling and confirms that assimilation of MODIS LAI data in watershed models can effectively improve both hydrology and water quality predictions.

Adnan Rajib↗

Research Data Alliance: Understanding Big Data Analytics Applications in Earth Science

The Research Data Alliance (RDA) enables data to be shared across barriers through focused working groups and interest groups, formed of experts from around the world - from academia, industry and government. Its Big Data Analytics (BDA) interest groups seeks to develop community based recommendations on feasible data analytics approaches to address scientific community needs of utilizing large quantities of data. BDA seeks to analyze different scientific domain applications (e.g. earth science use cases) and their potential use of various big data analytics techniques. These techniques reach from hardware deployment models up to various different algorithms (e.g. machine learning algorithms such as support vector machines for classification). A systematic classification of feasible combinations of analysis algorithms, analytical tools, data and resource characteristics and scientific queries will be covered in these recommendations. This contribution will outline initial parts of such a classification and recommendations in the specific context of the field of Earth Sciences. Given lessons learned and experiences are based on a survey of use cases and also providing insights in a few use cases in detail.

Riedel, Morris↗

Applications of Anomaly Detection and Precursor Identification in Airspace Operations

As we continue to advance the U.S. National Airspace into the next generation of air traffic, we face challenges in both increase in complexity, as well as, a significant growth in traffic volume. Addressing these challenges, while maintaining the same level of safety is an important application of data mining. Because of these significant shifts in airspace design and usage there is a need to identify current and emergent safety risks along with their potential precursors. In recent years NASA has made advancements in developing scalable methods to address this effort in the Big Data paradigm. Multiple kernel anomaly detection approaches have been employed on both surveillance radar data and flight operational quality assurance data to identify operationally significant safety risks. Additionally, events have been explored with a recently developed precursor identification tool to discover states that reveal an increased probability of a safety event. These tools can be used to discover emerging safety risks that may not be currently monitored, which allows for mitigation tactics to be employed and ultimately make the overall airspace safer. This talk will discuss an overview of these methods and a discussion of the findings.

anomaly detection↗

Macromolecules & Manufacturing Science

Outline • SRNL Overview • Mission overview • Polymers enabling the mission • R&D Highlights • Polymers in radiation environments • Tooling in shielded cells • Packaging for nuclear material shipments • Polymers supporting tank waste remediation • Ref electrode • Epoxy and polymer grout • Polymers for fusion energy • Deuterium labelling • Polymers for additive manufacturing • Coalescence and blends: experimental and predictive • Process modelling and sorting through big data (peregrine and latticeJ)

Chatham, Camden [Savannah River National Laborator↗

Mini-Stamp as a Micro-Display for At-A-Glance Subsystem Information for DSN Links

Operators of the Deep Space Network (DSN) attend to numerous tasks with the overall goal of providing continuous support for the world's deep space missions. This high-stakes operations environment requires operators to understand the state of the Deep Space Network and predict what will happen next. Under the Follow-the-Sun initiative which requires remote operations of the highly complex telecommunications equipment, operators will need to remain aware of the state of the entire network rather than just their own facility, and transitioning fluidly between periods of low activity and periods of high demand. I designed a micro-display for operators to see, at a glance, the state of a Deep Space Network support including its subsystems. Using in-depth user-centered and participatory design techniques to identify information requirements, I designed what I called a Postage Stamp (NTR-49720) for individual operators to be able to maintain awareness of their own assigned supports. However, under Follow the Sun, operators must remain aware of all supports. The area occupied by the Postage Stamp must shrink to allow operators to see the state of the entire system, e.g., via a Big Board posted prominently in the operations room. Micro-displays are tools for mental model re-alignment, helping operators to keep their mental models of how the system works and behaves aligned with the changing state of the complex system. Data-driven micro-displays such as the Postage Stamp and Mini-Stamp display information about the system in a consistent way. Like a traffic light, the format of the micro-display never changes: the operator always knows where to look to find a specific piece of information. The Mini-Stamp always looks like the Mini-Stamp, and all of its data fields always lie in the same place on the micro-display. Real-time data flows through the Mini-Stamp to provide information to the operator.

Holloway, Alexandra↗