Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data landscape”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Creating the First Interactive Mobility Data Landscape

As the world continues to be driven increasingly driven by data, the ways this data is sorted and collected is critical for researchers and specification creators. In the world of mobility data, there are few if any de jure standards connecting data, and there is little knowledge on the gaps that exist in the data. In this project we have created an interactive landscape where mobility data and data specifications can be categorized and organized in an easy-to-use living document.

ADVANCED PROPULSION SYSTEMS↗

Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

Machine learning (ML)-accelerated discovery requires large amounts of high-fidelity data to reveal predictive structure–property relationships. For many properties of interest in materials discovery, the challenging nature and high cost of data generation has resulted in a data landscape that is both scarcely populated and of dubious quality. Data-driven techniques starting to overcome these limitations include the use of consensus across functionals in density functional theory, the development of new functionals or accelerated electronic structure theories, and the detection of where computationally demanding methods are most necessary. When properties cannot be reliably simulated, large experimental data sets can be used to train ML models. In the absence of manual curation, increasingly sophisticated natural language processing and automated image analysis are making it possible to learn structure–property relationships from the literature. Finally, models trained on these data sets will improve as they incorporate community feedback.

36 MATERIALS SCIENCE↗

Generalized Relationship Linking Water Balance and Vegetation Productivity across Site-to-Regional Scales

Evapotranspiration (ET) is a pivotal component in catchment-scale water balance and is essential for informed watershed management. Nevertheless, uncertainties in ET observation or modeling have been hindering effective water resources management. This study addresses this gap by establishing a robust, generalized linear relationship between ET and gross primary productivity (GPP) at the catchment scale. We test the linearity of the relationships between monthly GPP and ET data at 380 near-natural catchments across various climatic and landscape conditions in the contiguous U.S., yielding Pearson’s r ≥ 0.6 for 97% of the 380 catchments. We then develop a regionalization strategy to parameterize this GPP-ET relationship at the catchment scale by identifying and utilizing the linkages between the parameter values and extensively available hydroclimatic and landscape data. We demonstrate the efficacy of the proposed GPP-ET relationship and parameter regionalization strategy by their combined predictive capacity, where the predicted monthly GPP matches well with remote-sensing-based GPP product, achieving Kling-Gupta Efficient (KGE) values ≥ 0.5 for 92% of the catchments. In addition, we verify the relationship and its parameter regionalization at 35 AmeriFlux sites with KGE ≥ 0.5 for 25 sites, suggesting that the new relationship is transferable across the site, catchment, and regional scales. Furthermore, our findings are valuable for improving remote-sensing-based estimation of monthly ET and diagnosing coupled water–carbon simulations in land surface and Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product↗

Multi-dimensional resilience: A quantitative exploration of disease outcomes and economic, political, and social resilience to the COVID-19 pandemic in six countries

The COVID-19 pandemic has highlighted a need for better understanding of countries’ vulnerability and resilience to not only pandemics but also disasters, climate change, and other systemic shocks. A comprehensive characterization of vulnerability can inform efforts to improve infrastructure and guide disaster response in the future. In this paper, we propose a data-driven framework for studying countries’ vulnerability and resilience to incident disasters across multiple dimensions of society. To illustrate this methodology, we leverage the rich data landscape surrounding the COVID-19 pandemic to characterize observed resilience for several countries (USA, Brazil, India, Sweden, New Zealand, and Israel) as measured by pandemic impacts across a variety of social, economic, and political domains. We also assess how observed responses and outcomes (i.e., resilience) of the COVID-19 pandemic are associated with pre-pandemic characteristics or vulnerabilities, including (1) prior risk for adverse pandemic outcomes due to population density and age and (2) the systems in place prior to the pandemic that may impact the ability to respond to the crisis, including health infrastructure and economic capacity. Our work demonstrates the importance of viewing vulnerability and resilience in a multi-dimensional way, where a country’s resources and outcomes related to vulnerability and resilience can differ dramatically across economic, political, and social domains. This work also highlights key gaps in our current understanding about vulnerability and resilience and a need for data-driven, context-specific assessments of disaster vulnerability in the future.

59 BASIC BIOLOGICAL SCIENCES↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

High-frequency Data Integration for Landscape Model Calibration of Carbon Fluxes Across Diverse Tidal Marshes

Terrestrial Aquatic Interfaces (TAIs), and tidal wetlands in particular, store large amounts of carbon yet are not well represented in Earth System Models (ESMs). Predictions of carbon cycling and greenhouse gas (GHG) emissions in tidal wetlands are highly uncertain. Eddy covariance (EC) towers provide ecosystem-scale GHG flux data at a temporal resolution (every 30min) that is helpful for parameterizing and improving mechanistic realism in ESMs. We propose to use a network of eddy covariance towers and standardized ancillary data streams, along with mesocosm experiments and statistical analyses, across diverse tidal wetlands of North America to develop and improve biogeochemical modeling at the TAI. Our overarching objective is to improve understanding and process-based modeling of gross primary productivity (GPP) and CH4 emission responses, both non-linear and asynchronous, to stressors including plant inundation, disturbance, salinity and nitrogen loading.

54 ENVIRONMENTAL SCIENCES↗

Landscape analysis of environmental data sources for linkage with SEER cancer patients database

Abstract One of the challenges associated with understanding environmental impacts on cancer risk and outcomes is estimating potential exposures of individuals diagnosed with cancer to adverse environmental conditions over the life course. Historically, this has been partly due to the lack of reliable measures of cancer patients’ potential environmental exposures before a cancer diagnosis. The emerging sources of cancer-related spatiotemporal environmental data and residential history information, coupled with novel technologies for data extraction and linkage, present an opportunity to integrate these data into the existing cancer surveillance data infrastructure, thereby facilitating more comprehensive assessment of cancer risk and outcomes. In this paper, we performed a landscape analysis of the available environmental data sources that could be linked to historical residential address information of cancer patients’ records collected by the National Cancer Institute’s Surveillance, Epidemiology, and End Results Program. The objective is to enable researchers to use these data to assess potential exposures at the time of cancer initiation through the time of diagnosis and even after diagnosis. The paper addresses the challenges associated with data collection and completeness at various spatial and temporal scales, as well as opportunities and directions for future research.

60 APPLIED LIFE SCIENCES↗

The future low-temperature geochemical data-scape as envisioned by the U.S. geochemical community

Data sharing benefits the researcher, the scientific community, and the public by allowing the impact of data to be generalized beyond one project and by making science more transparent. However, many scientific communities have not developed protocols or standards for publishing, citing, and versioning datasets. One community that lags in data management is that of low-temperature geochemistry (LTG). This paper resulted from an initiative from 2018 through 2020 to convene LTG and data scientists in the U.S. to strategize future management of LTG data. Through webinars, a workshop, a preprint, a townhall, and a community survey, the group of U.S. scientists discussed the landscape of data management for LTG – the data-scape. Currently this data-scape includes a “street bazaar” of data repositories. This was deemed appropriate in the same way that LTG scientists publish articles in many journals. The variety of data repositories and journals reflect that LTG scientists target many different scientific questions, produce data with extremely different structures and volumes, and utilize copious and complex metadata. Nonetheless, the group agreed that publication of LTG science must be accompanied by sharing of data in publicly accessible repositories, and, for sample-based data, registration of samples with globally unique persistent identifiers. LTG scientists should use certified data repositories that are either highly structured databases designed for specialized types of data, or unstructured generalized data systems. Recognizing the need for tools to enable search and cross-referencing across the proliferating data repositories, the group proposed that the overall data informatics paradigm in LTG should shift from “build data repository, data will come” to “publish data online, cybertools will find”. Funding agencies could also provide portals for LTG scientists to register funded projects and datasets, and forge approaches that cross national boundaries. Finally, the needed transformation of the LTG data culture requires emphasis in student education on science and management of data.

58 GEOSCIENCES↗

Accelerating Lossy and Lossless Compression on Emerging BlueField DPU Architectures

Data compression has become a crucial technique in addressing performance bottlenecks caused by increasing data volumes in High-Performance Computing (HPC), Big Data, and Deep Learning (DL). Despite its potential to boost system performance, recent studies have identified significant challenges with existing compression methods, mainly due to their high computational demands amidst continuously growing data sizes. Concurrently, the advent of Data Processing Units (DPUs), equipped with programmable System-on-Chip (SoC) and specialized compression accelerators, offers a promising opportunity to alter the landscape of data compression. This paper explores the complexities and potential of leveraging NVIDIA BlueField DPUs to accelerate lossy and lossless compression. Towards this, we introduce PEDAL, an innovative library that leverages the hardware capabilities of DPUs to unify and optimize data compression designs. Moreover, we seamlessly co-design PEDAL with the popular MPICH MPI library, demonstrating up to 101x speedup in compression time and 88x decrease in communication latency. Drawing on these achievements, we share our experience with various research communities about accelerating data compression on DPUs in communication-oriented HPC scenarios.

Li, Yuke↗

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗

Shedding light on U.S. small and midsize data centers: Exploring insights from the CBECS survey

As demand for digital services accelerates, the energy and environmental footprint of data centers faces increasing scrutiny. While hyperscale cloud facilities have driven efficiency gains, small and midsize U.S. data centers remain a critical yet underexamined segment with significant untapped potential for energy savings. This study leverages data from the Commercial Buildings Energy Consumption Survey (CBECS) to analyze trends in server stocks, computing customers, cooling system adoption and efficiency, and geospatial distribution from 2012 to 2018. Findings reveal a sharp decline in small and midsize data centers, from 1.764 million to 1.398 million, with server counts dropping from 5.177 million to 4.262 million—aligning with the broader shift toward cloud computing. More than 40 % of servers in small data centers and 55 % in midsize data centers are housed in office buildings, and over half of all servers are concentrated in climate zones 5A (cold), 3A (mixed-humid), and 4A (mixed-humid), with the highest densities in metropolitan hubs. While direct expansion units remain the dominant cooling system, a clear transition toward more energy-efficient solutions, particularly air economizers, is evident. By integrating server and cooling system distributions, we estimate Power Usage Effectiveness (PUE) and Water Usage Effectiveness (WUE) for U.S. data centers by size and year. Results show that midsize data centers are more energy-efficient but more water-intensive due to the widespread use of water-cooled chillers. These findings highlight the trade-offs in cooling system selection and provide a critical foundation for policies aimed at enhancing efficiency in an evolving data center landscape.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Habitat quality influences trade-offs in animal movement along the exploration–exploitation continuum

Abstract To successfully establish itself in a novel environment, an animal must make an inherent trade-off between knowledge accumulation and exploitation of knowledge gained (i.e., the exploration–exploitation dilemma). To evaluate how habitat quality affects the spatio-temporal scale of switching between exploration and exploitation during home range establishment, we conducted experimental trials comparing resource selection and space-use of translocated animals to those of reference individuals using reciprocal translocations between habitat types of differing quality. We selected wild pigs ( Sus scrofa ) as a model species to investigate hypotheses related to the movement behavior of translocated individuals because they are globally distributed large mammals that are often translocated within their introduced range to facilitate recreational hunting. Individuals translocated to higher quality habitat (i.e. higher proportions of bottomland hardwood habitats) exhibited smaller exploratory movements and began exploiting resources more quickly than those introduced to lower quality areas, although those in lower-quality areas demonstrated an increased rate of selection for preferred habitat as they gained knowledge of the landscape. Our data demonstrate that habitat quality mediates the spatial and temporal scale at which animals respond behaviorally to novel environments, and how these processes may determine the success of population establishment.

59 BASIC BIOLOGICAL SCIENCES↗

Disaster risk and artificial intelligence: A framework to characterize conceptual synergies and future opportunities

Artificial intelligence (AI) methods have revolutionized and redefined the landscape of data analysis in business, healthcare, and technology. These methods have innovated the applied mathematics, computer science, and engineering fields and are showing considerable potential for risk science, especially in the disaster risk domain. The disaster risk field has yet to define itself as a necessary application domain for AI implementation by defining how to responsibly balance AI and disaster risk. (1) How is AI being used for disaster risk applications; and how are these applications addressing the principles and assumptions of risk science, (2) What are the benefits of AI being used for risk applications; and what are the benefits of applying risk principles and assumptions for AI-based applications, (3) What are the synergies between AI and risk science applications, and (4) What are the characteristics of effective use of fundamental risk principles and assumptions for AI-based applications? This study develops and disseminates an online survey questionnaire that leverages expertise from risk and AI professionals to identify the most important characteristics related to AI and risk, then presents a framework for gauging how AI and disaster risk can be balanced. This study is the first to develop a classification system for applying risk principles for AI-based applications. This classification contributes to understanding of AI and risk by exploring how AI can be used to manage risk, how AI methods introduce new or additional risk, and whether fundamental risk principles and assumptions are sufficient for AI-based applications.

97 MATHEMATICS AND COMPUTING↗

Integrating AI with physics-based hydrological models and observations for insightinto changing climate and anthropogenic impacts

Focal Areas: Advanced computational methods that integrate AI, physics, and observations to provide predictive landscape hydrological modeling over large areas (regional, continental, worldwide) while incorporating increasingly available high-resolution data from drones, lidar and satellite. Science Challenge: Landscape data is available at finer scales than can be used in physics-based hydrological (PBH) models for regional or continental terrestrial water modeling. Thus, we throw away observable detail to achieve computability. We argue that integration of AI with PBH models and observed data can be used to provide upscaling for predictive models that are computable, retain physical conservation properties, and represent the fine-scale features that affect complex flow physics through both natural and urban environments. Developing such next-generation capabilities requires outside-the-box thinking that melds the different approaches of AI modeling, PBH modeling, and observation across multiple scales from local drones to satellites.

54 ENVIRONMENTAL SCIENCES↗

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie↗

Denudation, solute export, landscape evolution modeling, and geographic information system data for the East River watershed, Colorado, USA (2020-2024)

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of lithologic controls on denudation, solute export, carbon-scaling relationships, and transient landscape evolution in the East River watershed near Crested Butte, Colorado, USA. The package includes GIS layers used to produce the Figure 2 map, including drainage, hillshade, lithology, sample locations, and basin polygons, together with comma-separated value (CSV) tables and matching CSV data dictionaries. One group of tables reports sample-level and catchment-level information for river-sediment samples analyzed for in situ-produced cosmogenic beryllium-10 (10Be), including sample names, outlet elevations, geographic coordinates, upstream drainage area, rock-type classes, production-rate scaling scheme, analyzed nuclide, catchment-averaged denudation rates, and associated lower and upper analytical uncertainties. Sample and catchment attributes provide the basis for comparing denudation rates across intrusive, shale, sedimentary, and mixed-lithology settings. A second group of tables reports supporting information for landscape-evolution modeling and the mapped geologic framework of the study area. Included files list parameter values and definitions for the two-phase landscape-evolution simulations, summarize full-domain model erosion fluxes and topographic metrics for different simulation configurations, provide a fixed-area carbon-model scaling table, and summarize mapped geologic units within the East River study domain, including geologic code, formation name, lithologic description, mapped area, and lithologic class grouping. Model outputs and geologic summaries support interpretation of transient landscape behavior and its relation to the mapped distribution of shale, intrusive, sedimentary, and surficial units. A third group of tables reports hydrologic and hydrochemical information used to quantify dissolved export from the watershed. Included files provide site-level values for drainage area, mean annual solute export, standard error of annual export, area-normalized solute yield, and equivalent weathering rate for five East River monitoring sites, along with metadata describing the number, sampling cadence, and date range of discharge records and partial and full total dissolved solids observations used in the solute-yield analyses. The package also contains a supplementary daily ion-load time series with daily mean discharge, discharge observation counts, dissolved concentrations, and daily loads for calcium, magnesium, sodium, potassium, chloride, sulfate, nitrate, fluoride, dissolved silica, charge-balance bicarbonate, and total dissolved solids. The package contains GIS files, comma-separated value files (.csv), CSV data dictionaries, a file-level metadata table, a package-tree text file, and a readme text file.

10Be↗

Stable Isotope and Geochemical Evidence for Hydrological Isolation in an Arctic Coastal Plain Landscape, Barrow, Alaska, 2013

Data include results from water chemistry and water isotope analyses for samples collected in Barrow, Alaska during July and September 2013. Samples were from surface and soil pore waters from 15 locations: 3 locations from interlake polygonal terrain, 6 locations associated with interlake drainages, and 6 locations within or at the outlets of different aged drained thaw lake basins (DTLBs). Samples were taken in different drainage flow types at three different depths at each location in and around the Barrow Environmental Observatory. This dataset includes one .csv data file and one .pdf user guide.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗