Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

J-PLUS: Support vector machine applied to STAR-GALAXY-QSO classification

Context. In modern astronomy, machine learning has proved to be efficient and effective in mining big data from the newest telescopes. Aims. In this study, we construct a supervised machine-learning algorithm to classify the objects in the Javalambre Photometric Local Universe Survey first data release (J-PLUS DR1). Methods. The sample set is featured with 12-waveband photometry and labeled with spectrum-based catalogs, including Sloan Digital Sky Survey spectroscopic data, the Large Sky Area Multi-Object Fiber Spectroscopic Telescope, and VERONCAT – the Veron Catalog of Quasars & AGN. The performance of the classifier is presented with the applications of blind test validations based on RAdial Velocity Extension, the Kepler Input Catalog, the Two Micron All Sky Survey Redshift Survey, and the UV-bright Quasar Survey. A new algorithm was applied to constrain the potential extrapolation that could decrease the performance of the machine-learning classifier. Results. The accuracies of the classifier are 96.5% in the blind test and 97.0% in training cross-validation. The F1-scores for each class are presented to show the balance between the precision and the recall of the classifier. We also discuss different methods to constrain the potential extrapolation.

79 ASTRONOMY AND ASTROPHYSICS↗

Towards Automated Analytics of Research Publications

For readers of scientific publications it remains a big challenge to unambiguously relate the published research with the data used. To a substantial degree it is attributed to authors, journals, editors, and reviewers not prioritizing correct data citation, which impacts traceability, repeatability, and giving credits to published authors and their funding sources. Furthermore, uniform classification of the content of the published research is hampered by journals using journal specific topics and letting authors to assign free text keywords to their papers. We demonstrate automated analytics methods for extracting and relating datasets used and the research application areas by processing 1,300 research papers that referenced the NASA Giovanni service (but probably not the datasets in particular) as supporting their publication process. This presentation was given during the 2022 ESIP January meeting held virtually in January 2022.

Irina Gerasimov↗

NASA Soil Moisture Active Passive Mission Status and Science Highlights

The Soil Moisture Active Passive (SMAP) observatory was launched January 31, 2015, and its L-band radiometer and radar instruments became operational during April 2015. This paper provides a summary of the quality assessment of its baseline soil moisture and freeze/thaw products as well as an overview of new products. The first new product explores the Backus Gilbert optimum interpolation based on the oversampling characteristics of the SMAP radiometer. The second one investigates the disaggregation of the SMAP radiometer data using the European Space Agency's Sentinel-1 C-band synthetic aperture radar (SAR) data to obtain soil moisture products at about 1 to 3 km resolution. In addition, SMAPs L-band data have been found useful for many scientific applications, including depictions of water cycles, vegetation opacity, ocean surface salinity and hurricane ocean surface wind mapping. Highlights of these new applications will be provided.The SMAP soil moisture, freeze/taw state and SSSprovide a synergistic view of water cycle. For example, Fig.7 illustrates the transition of freeze/thaw state, change of soilmoisture near the pole and SSS in the Arctic Ocean fromApril to October in 2015 and 2016. In April, most parts ofAlaska, Canada, and Siberia remained frozen. Melt onsetstarted in May. Alaska, Canada, and a big part of Siberia havebecome thawed at the end of May; some freshwater dischargecould be found near the mouth of Mackenzie in 2016, but notin 2015. The soil moisture appeared to be higher in the Oband Yenisei river basins in Siberia in 2015. As a result,freshwater discharge was more widespread in the Kara Seanear the mouths of both rivers in June 2015 than in 2016. TheNorth America and Siberia have become completely thawedin July. After June, the freshwater discharge from other riversinto the Arctic, indicated by blue, also became visible. Thefreeze-up started in September and the high latitude regionsin North America and Eurasia became frozen. Comparing thespread of freshwater in August 2015 and 2016 suggests thatthere was more discharge from Ob and Yenisei in 2015,which appeared to correspond to a higher soil moisturecontent in the Ob and Yenisei basins. In contrast, Mackenzieappeared to have more discharge in September 2016.

soil moisture↗

Challenges and Opportunities in Magnetospheric Space Weather Prediction

Abstract Space Weather is the study of the dynamics of the coupled Solar‐Terrestrial environment, as these dynamics impact technological systems and human activity. This paper reviews a selection of the advances, challenges, and new opportunities for magnetospheric space weather, and while the focus is on specific phenomena, many other aspects of space weather will have similar challenges and needs. As a scientific field with direct applications, the field of space weather is partly driven by imperatives from both policy and operations. We provide an introduction to some of the context in which the field exists and discuss how this might shape future developments and norms within the space weather enterprise. We briefly examine benchmarking, as a policy and operationally driven activity, as it provides immediate societal relevance and an opportunity to stretch scientific understanding. As numerical space weather prediction now becomes routine, and exascale computing is in the near future, we identify challenges relating to computational expense and big data, capturing and accounting for uncertainties, and specification of boundary conditions. Here, as with the observations supporting numerical space weather prediction, the key challenge lies in extending the lead time of predictions. We also discuss the role of data, particularly in regard to model validation and empirical modeling. Due to the growing societal impact of space weather, we also examine the relationships between space weather and its terrestrial counterpart and look at the importance of continuous evaluation, monitoring progress in predictive capability, and communication with researchers, forecasters, and end users.

79 ASTRONOMY AND ASTROPHYSICS↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Advanced Health Information Technology Analytic Framework and Application to Hazard Detection

Health Information Technology (HIT) aims to improve healthcare outcomes by organizing and analyzing various health-related data. With data accumulating at a staggering rate, the importance of real-time analytics has been increasing dramatically, shifting the focus of informatics from batch processing to streaming analytics. HIT is also facing unprecedented challenges in adapting to this new requirement and leveraging advanced IT technologies. This paper introduces a HIT data and compute platform that supports multi-granularity real-time analytics from heterogeneous data sources. The paper first identifies functional requirements and proposes a framework that satisfies the requirements using state-of-the-art big data technologies including Apache Kafka, Spark Structured Streaming Engine, and Delta Lake. To demonstrate its capability to support data analytics in multiple time granularities analytics, a statistical process control-based hazard detection algorithm has been implemented on top of the framework to detect unexpected hazards from order cancellation data of the Department of US Veterans Affairs (VA) in near real-time.

Kumar, Mohit↗

Data Mining and Machine Learning for Power System Monitoring, Understanding, and Impact Evaluation

This chapter presents results from the Big Data analysis framework to improve power system situational awareness and system reliability. For this purpose, a dataset with real-world phasor measurement unit data and historical transmission system outage data has been created and used to carry out the analysis. Several statistical analysis and machine learning methods have been developed and implemented for event and anomaly detection and modeling. Detection and analysis results for actual examples of power system events are presented. Finally, data-driven characterization and risk assessment methods for weather-related extremes in power systems are developed and demonstrated on the Bonneville Power Administration system. These applications demonstrate the capability of Machine Learning (ML) methods to monitor system abnormalities, to predict system events, and to characterize the impact of extreme events on power grid

data mining, power grid, machine learning, anomaly↗

Aerial Captured Data and Processed Models in Beaumont-Port Arthur Region in Feb and Oct, 2023

Our Co-design team is from the University of Texas, working on a Department of Energy-funded project focused on the Beaumont-Port Arthur area. As part of this project, we will be developing climate-resilient design solutions for areas of the region. More on www.caee.utexas.edu.We used a DJI Mavic 2 Pro to capture aerial photos in Beaumont-Port Arthur, TX, in February 2023, including:I. Beaumont Soccer ClubII. Corps’ Port Arthur Resident OfficeIII. Halbouty Pump Station comprises its vicinityIV. Lamar University (Including Exxon Power Plants close to Lamar Univ.)V. MLK Boulevard for aerial images of the industry and the ship channelVI. Salt Water Barrier (include some aerial images about the Big Thicket)Aerial photos taken were through DroneDeploy autonomous flight, and models were processed through the DroneDeploy engine as well. All aerial photos are in .JPG format and contained in zipped files for each location.The processed data package including 3D models, geospatial data, mappings, point clouds, and the animation video of Halbouty Pump Station has various file types:- The Adobe Suite gives you great software to open .Tif files.- You can use LASUtility (Windows), ESRI ArcGIS Pro (Windows), or Blaze3D (Windows, Linux) to open a LAS file and view the data it contains.- Open an .OBJ file with a large number of free and commercial applications. Some examples include Microsoft 3D Builder, Apple Preview, Blender, and Autodesk.- You may use ArcGIS, Merkaartor, Blender (with the Google Earth Importer plug-in), Global Mapper, and Marble to open .KML files.- The .tfw world file is a text file used to georeference the GeoTIFF raster images, like the orthomosaic and the DSM. You need suitable software like ArcView to open a .TFW file.This dataset provides researchers with sufficient geometric data and the status quo of the land surface at the locations mentioned above. This dataset could streamline researchers' decision-making processes and enhance the design as well.In October 2023, we had our follow-up data collection, including:I. Beaumont Soccer ClubII. Shipping and Receiving Center at Lamar UniversityAfter the aerial collection, we obtained aerial photos of those two locations mentioned above, as well as processed data (such as point clouds and models).

2D mapping↗

Where Is the Provenance? Ethical Replicability and Reproducibility in GIScience and Its Critical Applications

As replicability and reproducibility (R&R) crises develop within emerging convergent inquiry, ethical use of provenance information is central to the establishment and preservation of trust in critical applications of GIScience and geospatial technologies. Today large volumes of geospatial data are generated at high velocity from satellite sensors and unmanned aircraft systems, citizen sensors, geolocation-based data services, global navigation satellite systems, and so on. The extensive use of these data for applications such as disaster and humanitarian response raises the issue of R&R from competing perspectives of location privacy and geospatial data quality. Although geospatial data can be integrated and linked with contextual information to identify individuals’ movements, steps taken to ensure privacy can complicate the multiuser development of high-quality geospatial workflows. Provenance information as digital records of historical (retrospective) and potential future (prospective) geospatial processes is often overlooked, misunderstood, or inadequately addressed. We explore the relationship between provenance information, location privacy, and geospatial data quality in the context of R&R with a focus on disaster analytics. Here, we argue that in the era of big data and deep learning, GIScientists and associated institutions bear greater responsibility both for geospatial workflow quality and for location privacy. Given vastly heterogenous computational landscapes, we provide practical recommendations for ethically driven provenance and R&R research and development within the GIScience community and beyond.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Furthermore, our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.

97 MATHEMATICS AND COMPUTING↗

Proceedings of the 39th Lunar and Planetary Science Conference

Sessions with oral presentations include: A SPECIAL SESSION: MESSENGER at Mercury, Mars: Pingos, Polygons, and Other Puzzles, Solar Wind and Genesis: Measurements and Interpretation, Asteroids, Comets, and Small Bodies, Mars: Ice On the Ground and In the Ground, SPECIAL SESSION: Results from Kaguya (SELENE) Mission to the Moon, Outer Planet Satellites: Not Titan, Not Enceladus, SPECIAL SESSION: Lunar Science: Past, Present, and Future, Mars: North Pole, South Pole - Structure and Evolution, Refractory Inclusions, Impact Events: Modeling, Experiments, and Observations, Mars Sedimentary Processes from Victoria Crater to the Columbia Hills, Formation and Alteration of Carbonaceous Chondrites, New Achondrite GRA 06128/GRA 06129 - Origins Unknown, The Science Behind Lunar Missions, Mars Volcanics and Tectonics, From Dust to Planets (Planetary Formation and Planetesimals):When, Where, and Kaboom! Astrobiology: Biosignatures, Impacts, Habitability, Excavating a Comet, Mars Interior Dynamics to Exterior Impacts, Achondrites, Lunar Remote Sensing, Mars Aeolian Processes and Gully Formation Mechanisms, Solar Nebula Shake and Bake: Mixing and Isotopes, Lunar Geophysics, Meteorites from Mars: Shergottite and Nakhlite Invasion, Mars Fluvial Geomorphology, Chondrules and Chondrule Formation, Lunar Samples: Chronology, Geochemistry, and Petrology, Enceladus, Venus: Resurfacing and Topography (with Pancakes!), Overview of the Lunar Reconnaissance Orbiter Mission, Mars Sulfates, Phyllosilicates, and Their Aqueous Sources, Ordinary and Enstatite Chondrites, Impact Calibration and Effects, Comparative Planetology, Analogs: Environments and Materials, Mars: The Orbital View of Sediments and Aqueous Mineralogy, Planetary Differentiation, Titan, Presolar Grains: Still More Isotopes Out of This World, Poster sessions include: Education and Public Outreach Programs, Early Solar System and Planet Formation, Solar Wind and Genesis, Asteroids, Comets, and Small Bodies, Carbonaceous Chondrites, Chondrules and Chondrule Formation, Chondrites, Refractory Inclusions, Organics in Chondrites, Meteorites: Techniques, Experiments, and Physical Properties, MESSENGER and Mercury, Lunar Science Present: Kaguya (SELENE) Results, Lunar Remote Sensing: Basins and Mapping of Geology and Geochemistry, Lunar Science: Dust and Ice, Lunar Science: Missions and Planning, Mars: Layered, Icy, and Polygonal, Mars Stratigraphy and Sedimentology, Mars (Peri)Glacial, Mars Polar (and Vast), Mars, You are Here: Landing Sites and Imagery, Mars Volcanics and Magmas, Mars Atmosphere, Impact Events: Modeling, Experiments, and Observation, Ice is Nice: Mostly Outer Planet Satellites, Galilean Satellites, The Big Giant Planets, Astrobiology, In Situ Instrumentation, Rocket Scientist's Toolbox: Mission Science and Operations, Spacecraft Missions, Presolar Grains, Micrometeorites, Condensation-Evaporation: Stardust Ties, Comet Dust, Comparative Planetology, Planetary Differentiation, Lunar Meteorites, Nonchondritic Meteorites, Martian Meteorites, Apollo Samples and Lunar Interior, Lunar Geophysics, Lunar Science: Geophysics, Surface Science, and Extralunar Components, Mars, Remotely, Mars Orbital Data - Methods and Interpretation, Mars Tectonics and Dynamics, Mars Craters: Tiny to Humongous, Mars Sedimentary Mineralogy, Martian Gullies and Slope Streaks, Mars Fluvial Geomorphology, Mars Aeolian Processes, Mars Data and Mission,s Venus Mapping, Modeling, and Data Analysis, Titan, Icy Dwarf Satellites, Rocket Scientist's Toolbox: In Situ Analysis, Remote Sensing Approaches, Advances, and Applications, Analogs: Sulfates - Earth and Lab to Mars, Analogs: Remote Sensing and Spectroscopy, Analogs: Methods and Instruments, Analogs: Weird Places!. Print Only Early Solar System, Solar Wind, IDPs, Presolar/Solar Grains, Stardust, Comets, Asteroids, and Phobos, Venus, Mercury, Moon, Meteorites, Mars, Astrobiology, Impacts, Outer Planets, Satellites, and Rings, Support for Mission Operations, Analog Education and Public Outreach.

Source record↗

Big Bend Ecological Forecasting: Integrating Earth Observations Into Invasive Species Management Decisions in Big Bend National Park in Texas

Located in Brewster County Texas, U.S., along the Texas-Mexico border, Big Bend National Park is 3,243 square kilometers of desert, mountains, and rivers. NASA’s MSFC Spring 2024 DEVELOP Team partnered with the National Park Service (NPS), to address the environmental concern of perineal invasive grasses in Big Bend National Park. Buffelgrass (Cenchrus ciliaris), introduced to the park in the 1940’s, poses an ongoing threat, and causes habitat destruction for many of the park’s native ecosystems. Buffelgrass amplifies fire risk in the park, aids in the destruction of historic structures, and alters stream channels. To address the rising concern of Buffelgrass presence, unique advanced spatial technique applications were necessary to construct a habitat suitability model and perform a comprehensive fire risk assessment. A habitat suitability model was developed considering climate, vegetation, and phenological variables, in addition to physical and topographical variables. Subsequently, a fire risk model was developed, taking into account fire history data, accessibility factors, climate trends and predictions, along with developed areas. Multi-Source Land Surface Phenology (LSP) Sentinel-2 and Landsat 8 Operational Land Imager (OLI) imagery were used to predict Buffelgrass hotspot locations throughout the park. These analyses allowed the identification of optimal Buffelgrass habitat and hotspot locations, as well as park zones that reflect the greatest risk for future Buffelgrass invasion and fire risk. The collective results of the habitat suitability model, fire risk assessment, and Buffelgrass hot spot identification will allow the NPS to facilitate efficient mitigation measures and management strategies, and improved resource allocation where it’s most needed.

remote sensing↗

Physics-informed machine learning

Despite great progress in simulating multiphysics problems using the numerical discretization of partial differential equations (PDEs), one still cannot seamlessly incorporate noisy data into existing algorithms, mesh generation remains complex, and high-dimensional problems governed by parameterized PDEs cannot be tackled. Moreover, solving inverse problems with hidden physics is often prohibitively expensive and requires different formulations and elaborate computer codes. Machine learning has emerged as a promising alternative, but training deep neural networks requires big data, not always available for scientific problems. Instead, such networks can be trained from additional information obtained by enforcing the physical laws (for example, at random points in the continuous space-time domain). Such physics-informed learning integrates (noisy) data and mathematical models, and implements them through neural networks or other kernel-based regression networks. Moreover, it may be possible to design specialized network architectures that automatically satisfy some of the physical invariants for better accuracy, faster training and improved generalization. Furthermore, we review some of the prevailing trends in embedding physics into machine learning, present some of the current capabilities and limitations and discuss diverse applications of physics-informed learning both for forward and inverse problems, including discovering hidden physics and tackling high-dimensional problems.

97 MATHEMATICS AND COMPUTING↗

Evaluating E3SM Global Storm‐Resolving Model Simulations of Deep Convection: Insights From DP‐SCREAM During TRACER

Global Storm-Resolving Models (GSRMs) are becoming increasingly vital for advancing climate modeling and improving the prediction of extreme weather events. Houston, a coastal region frequently affected by deep convective storms, offers an ideal setting to evaluate the ability of GSRMs to simulate deep convection. This study assesses the performance of the Doubly Periodic Simple Cloud-Resolving E3SM (Energy Exascale Earth System Model) Atmosphere Model (DP-SCREAM) using observations from the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign. DP-SCREAM effectively reproduces the diurnal cycles of clouds and precipitation, demonstrating much greater skill than the E3SM single column model. The DP-SCREAM is demonstrated to be applicable to coastal regions, partially due to the forcing data sets already capturing the influence of breezes. DP-SCREAM also replicates biases persistent in the global version of SCREAM: the underrepresentation of boundary layer shallow clouds, a lack of mid-level congestus clouds, and the popcorn convection, characterized by small and disorganized convective cells generating the strongest precipitation. To investigate these issues, two sensitivity experiments were conducted: increasing the mixing length and scaling up the buoyancy flux within the Simplified Higher Order Closure scheme. Increasing the mixing length improved mid-level congestus representation and reduced unrealistic early morning fog occurrence. Enhancing buoyancy flux only marginally improved the bias of underproduced big convective cells. In conclusion, an additional resolution sensitivity test at 0.5 km grid spacing demonstrated that a refined horizontal resolution alone is insufficient to resolve these biases.

54 ENVIRONMENTAL SCIENCES↗

Enabling discovery data science through cross-facility workflows

Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.

Antypas, Katerina B.↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Multi-Branch Decoder Network Approach to Adaptive Temporal Data Selection and Reconstruction for Big Scientific Simulation Data

A key challenge in scientific simulation is that the simulation outputs often require intensive I/O and storage space to store the results for effective post hoc analysis. This article focuses on a quality-aware adaptive temporal data selection and reconstruction problem where the goal is to adaptively select simulation data samples at certain key timesteps in situ and reconstruct the discarded samples with quality assurance during post hoc analysis. This problem is motivated by the limitation of current solutions that a significant amount of simulation data samples are either discarded or aggregated during the sampling process, leading to inaccurate modeling of the simulated phenomena. Two unique challenges exist: 1) the sampling decisions have to be made in situ and adapted to the dynamics of the complex scientific simulation data; 2) the reconstruction error must be strictly bounded to meet the application requirement. To address the above challenges, we develop DeepSample , an error-controlled convolutional neural network framework, that jointly integrates a set of coherent multi-branch deep decoders to effectively reconstruct the simulation data with rigorous quality assurance. The results on two real-world scientific simulation applications show that DeepSample significantly outperforms other state-of-the-art methods on both sampling efficiency and reconstructed simulation data quality.

Zhang, Yang↗

Combining Large Datasets - Cancer Moonshot Task Group Final Summary

In February 2022, President Biden re-ignited the Cancer Moonshot with bold new goals: to reduce the cancer death rate by half within 25 years and improve the lives of people with cancer and cancer survivors. To achieve these ambitious goals, the White House convened the first-ever Cancer Cabinet, bringing together departments and agencies from across the federal government to end cancer as we know it.The Cancer Cabinet convened three task forces and supporting task groups, including the Data and Innovation Task Force, which supported the Cancer Moonshot priority to “Deliver innovation to patients and communities.” In early 2023, the “Combining Large Datasets” (CoLD) Task Group was created within the Data and Innovation Task Force. The scope of the CoLD Task Group was how federal agencies combine large datasets for broad applications across cancer prevention and control, including nutrition, epidemiology, and military/Veteran health. Within this scope, the group sought to better leverage the immense potential of data and power of data tools to increase our understanding of cancer incidence, causes, mortality, treatments, prevention, outcomes, costs, and all other aspects of the burden of cancer.

data integration↗