AI Surrogate Model for Distributed Computing Workloads
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.
Major changes are taking place in the way astronomy gets done. There are continuing advances in observational capabilities across the frequency spectrum, involving both ground-based and space-based facilities. There is also very rapid evolution of relevant computing and data management technologies. However, although the new technologies are filtering in to the astronomy community, and astronomers are looking at their computing needs in new ways, there is little coordination or coherent policy. Furthermore, although there is great awareness of the evolving technologies in the arena of operations, much of the existing operations infrastructure is ill-suited to take advantage of them. Astronomy, especially space astronomy, has often been at the cutting edge of computer use in data reduction and image analysis, but has been somewhat removed from advanced applications in operations, which have tended to be implemented by industry rather than by the end-user scientists. The purpose of this paper is threefold. First, we briefly review the background and general status of astronomy-related computing. Second, we make recommendations in three areas: data analysis; operations (directed primarily to NASA-related activities); and issues of management and policy, believing that these must be addressed to enable technological progress and to proceed through the next decade. Finally, we recommend specific NASA-related work as part of the Astrotech-21 plans, to enable better science operations in the operations of the Great Observatories and in the lunar outpost era.
Progress is reported on the development of SCOTTY, an expert knowledge-based system to automate the analysis procedure following test firings of the Space Shuttle Main Engine (SSME). The integration of a large-scale relational data base system, a computer graphics interface for experts and end-user engineers, potential extension of the system to flight engines, application of the system for training of newly-hired engineers, technology transfer to other engines, and the essential qualities of good software engineering practices for building expert knowledge-based systems are among the topics discussed.
We describe the use of scenarios to develop and refine requirement tables for parts of the Earth Observing System Data and Information System (EOSDIS). The National Aeronautics and Space Administration (NASA) is developing EOSDIS as part of its Mission-To-Planet-Earth (MTPE) project to accept instrument/platform observation requests from end-user scientists, schedule and perform requested observations of the Earth from space, collect and process the observed data, and distribute data to scientists and archives. Current requirements for the system are managed with tools that allow developers to trace the relationships between requirements and other development artifacts, including other requirements. In addition, the user community (e.g., earth and atmospheric scientists), in conjunction with NASA, has generated scenarios describing the actions of EOSDIS subsystems in response to user requests and other system activities. As part of a research effort in verification and validation techniques, this paper describes our efforts to develop requirements tables from these scenarios for the EOSDIS Core System (ECS). The tables specify event-driven mode transitions based on techniques developed by the Naval Research Lab's (NRL) Software Cost Reduction (SCR) project. The SCR approach has proven effective in specifying requirements for large systems in an unambiguous, terse format that enhance identification of incomplete and inconsistent requirements. We describe development of SCR tables from user scenarios and identify the strengths and weaknesses of our approach in contrast to the requirements tracing approach. We also evaluate the capabilities of both approach to respond to the volatility of requirements in large, complex systems.
Scientific experiments and computations, especially in High Energy Physics, are generating and accumulating data at an unprecedented rate. Effectively managing this vast volume of data while ensuring efficient data analysis poses a significant challenge for data centers, which must integrate various storage technologies. This paper proposes addressing this challenge by designing and developing a precise data popularity prediction model utilizing state-of-theart AI/ML techniques. This model is crafted from the analysis of ATLAS data and access patterns. It enables us to migrate infrequently accessed data to more economical storage media, such as tape drives, while storing frequently accessed data on faster yet costlier storage media like HDD or SSD. This strategic approach ensures data is placed optimally into the appropriate storage classes, thereby maximizing storage capacity while minimizing data access latency for end-users. Furthermore, the paper includes a performance evaluation of the prediction model using various key metrics such as F1 score, accuracy, precision and recall. Finally, we present a prototype use case, leveraging real-world file access data to assess the model’s impact on performance.
DATASTAR, Inc., of Picayune, Mississippi, has taken NASA s award-winning Earth Resources Laboratory Applications Software (ELAS) program and evolved it into a user-friendly desktop application and Internet service to perform processing, analysis, and manipulation of remotely sensed imagery data. NASA s Stennis Space Center developed ELAS in the early 1980s to process satellite and airborne sensor imagery data of the Earth s surface into readable and accessible information. Since then, ELAS information has been applied worldwide to determine soil content, rainfall levels, and numerous other variances of topographical information. However, end-users customarily had to depend on scientific or computer experts to provide the results, because the imaging processing system was intricate and labor intensive.
Since its inception over a decade ago, the DEVELOP National Program has provided students with experience in utilizing and integrating satellite remote sensing data into real world-applications. In 1998, DEVELOP began with three students and has evolved into a nationwide internship program with over 200 students participating each year. DEVELOP is a NASA Applied Sciences training and development program extending NASA Earth science research and technology to society. Part of the NASA Science Mission Directorate s Earth Science Division, the Applied Sciences Program focuses on bridging the gap between NASA technology and the public by conducting projects that innovatively use NASA Earth science resources to research environmental issues. Project outcomes focus on assisting communities to better understand environmental change over time. This is accomplished through research with global, national, and regional partners to identify the widest array of practical uses of NASA data. DEVELOP students conduct research in areas that examine how NASA science can better serve society. Projects focus on practical applications of NASA s Earth science research results. Each project is designed to address at least one of the Applied Sciences focus areas, use NASA s Earth observation sources and meet partners needs. DEVELOP research teams partner with end-users and organizations who use project results for policy analysis and decision support, thereby extending the benefits of NASA science and technology to the public.
The problems of managing and searching large archives of scientific journal articles can potentially be addressed through data mining and statistical techniques matured primarily for quantitative scientific data analysis. A journal paper could be represented by a multivariate descriptor, e.g., the occurrence counts of a number key technical terms or phrases (keywords), perhaps derived from a controlled vocabulary ( e . g . , the American Meteorological Society's Glossary of Meteorology) or bootstrapped from the journal archive itself. With this technique, conventional statistical classification tools can be leveraged to address challenges faced by both scientists and professional societies in knowledge management. For example, cluster analyses can be used to find bundles of "most-related" papers, and address the issue of journal bifurcation (when is a new journal necessary, and what topics should it encompass). Similarly, neural networks can be trained to predict the optimal journal (within a society's collection) in which a newly submitted paper should be published. Comparable techniques could enable very powerful end-user tools for journal searches, all premised on the view of a paper as a data point in a multidimensional descriptor space, e.g.: "find papers most similar to the one I am reading", "build a personalized subscription service, based on the content of the papers I am interested in, rather than preselected keywords", "find suitable reviewers, based on the content of their own published works", etc. Such services may represent the next "quantum leap" beyond the rudimentary search interfaces currently provided to end-users, as well as a compelling value-added component needed to bridge the print-to-digital-medium gap, and help stabilize professional societies' revenue stream during the print-to-digital transition.
The subject of sensor-based structural health monitoring is very diverse and encompasses a wide range of activities including initiatives and innovations involving the development of advanced sensor, signal processing, data analysis, and actuation and control technologies. In addition, it embraces the consideration of the availability of low-cost, high-quality contributing technologies, computational utilities, and hardware and software resources that enable the operational realization of robust health monitoring technologies. This report presents a detailed analysis of the cost benefit and other logistics and operational considerations associated with the implementation and utilization of sensor-based technologies for use in aerospace structure health monitoring. The scope of this volume is to assess the economic impact, from an end-user perspective, implementation health monitoring technologies on three structures. It specifically focuses on evaluating the impact on maintaining and supporting these structures with and without health monitoring capability.
The Mobile Bay region has experienced noteworthy land use and land cover (LULC) change in the latter half of the 20th century. Accompanying this change has been urban expansion and a reduction of rural land uses. Much of this LULC change has reportedly occurred since the landfall of Hurricane Frederic in 1979. The Mobile Bay region provides great economic and ecologic benefits to the Nation, including important coastal habitat for a broad diversity of fisheries and wildlife. Regional urbanization threatens the estuary s water quality and aquatic-habitat dependent biota, including commercial fisheries and avian wildlife. Coastal conservation and urban land use planners require additional information on historical LULC change to support coastal habitat restoration and resiliency management efforts. This presentation discusses results of a Gulf of Mexico Application Pilot project that was conducted in 2008 to quantify and assess LULC change from 1974 to 2008. This project was led by NASA Stennis Space Center and involved multiple Gulf of Mexico Alliance (GOMA) partners, including the Mobile Bay National Estuary Program (NEP), the U.S. Army Corps of Engineers, the National Oceanic and Atmospheric Administration s (NOAA s) National Coastal Data Development Center (NCDDC), and the NOAA Coastal Services Center. Nine Landsat images were employed to compute LULC products because of their availability and suitability for the application. The project also used Landsat-based national LULC products, including coastal LULC products from NOAA s Coastal Change & Analysis Program (C-CAP), available at 5-year intervals since 1995. Our study was initiated in part because C-CAP LULC products were not available to assess the region s urbanization prior to 1995 and subsequent to post Hurricane Katrina in 2006. This project assessed LULC change across the 34-year time frame and at decadal and middecadal scales. The study area included the majority of Mobile and Baldwin counties that encompass Mobile Bay. In doing so, each date of Landsat data was classified using an end-user defined modified Anderson level 1 classification scheme. LULC classifications were refined using a decision rule approach in conjunction with available C-CAP products. Individual dates of LULC classifications were validated by image interpretation of stratified random locations on raw Landsat color composite imagery in combination with higher resolution remote sensing and in-situ reference data. The results indicate that during the 34-year study period, urban areas increased from 96,688 to 150,227 acres, representing a 55.37% increase, or 1.63% per annum. Most of the identified urban expansion results from conversion of rural forest and agriculture to urban cover types. Final LULC mapping and metadata products were produced for the entire study area as well as watersheds of concern within the study area. Final project products, including LULC trend information, were incorporated into the Mobile Bay NEP State of the Bay report. Products and metadata were transferred to NOAA NCDDC to allow free online accessibility and use by GOMA partners and by the public.
The US Agency for International Development (USAID) s Famine Early Warning System Network (FEWS NET) provides monitoring and early warning support to decision makers responsible for responding to food insecurity emergencies on three continents. FEWS NET uses satellite remote sensing and ground observations of rainfall and vegetation in order to provide information on drought, floods and other extreme weather events to decision makers. Previous research has presented results from a professional review questionnaire with FEWS NET expert end-users whose focus was to elicit Earth observation requirements. The review provided FEWS NET operational requirements and assessed the usefulness of additional remote sensing data. Here we analyzed 1342 food security update reports from FEWS NET. The reports consider the biophysical, socioeconomic, and contextual influences on the food security in 17 countries in Africa from 2000-2009. The objective was to evaluate the use of remote sensing information in comparison with other important factors in the evaluation of food security crises. The results show that all 17 countries use rainfall information, agricultural production statistics, food prices and food access parameters in their analysis of food security problems. The reports display large scale patterns that are strongly related to history of the FEWS NET program in each country. We found that rainfall data was used 84% of the time, remote sensing of vegetation 28% of the time, and gridded crop models 10%, reflecting the length of use of each product in the regions. More investment is needed in training personnel on remote sensing products to improve use of data products throughout the FEWS NET system.
Here, in recent work [Vavrek et al. (2025)], we developed the performance optimization framework spectre-ml for gamma spectrometers with variable performance across many readout channels. The framework uses non-negative matrix factorization (NMF) and clustering to learn groups of similarly-performing channels and sweep through various learned channel combinations to optimize the performance tradeoff of including worse-performing channels for better total efficiency. In this work, we integrate the pyGEM uranium enrichment assay code with our spectre-ml framework, and show that the U-235 enrichment relative uncertainty can be directly used as an optimization target. We find that this optimization reduces relative uncertainties after a 30 -minute measurement by an average of 20%, as tested on six different H3D M400 CdZnTe spectrometers, which can significantly improve uranium non-destructive assay measurement times in nuclear safeguards contexts. Additionally, this work demonstrates that the spect re-ml optimization framework can accommodate arbitrary end-user spectroscopic analysis code and performance metrics, enabling future optimizations for complex Pu spectra.
Geothermal district energy systems (DES) with ambient-temperature loops, also known as thermal energy networks, are one option for decarbonizing space heating and cooling loads. Geothermal fifth-generation DES include an "ambient" temperature thermal loop that connects heat pumps at each building with thermal balancing sources such as geothermal borehole fields. Heating and cooling are provided via a water-source heat pump at each end-user. This project seeks to analyze the nationwide potential for ambient-temperature loop districts by creating a new module within the Distributed Geothermal Market Demand Model (dGeo). dGeo is an agent-based modeling tool for distributed geothermal resources; it can investigate potential on a nationwide or statewide scale using geospatial data for all 50 states and thermal demands for existing buildings. This process allows for high-level estimates of technical and economic potential for ambient-temperature loop districts across the United States. A lookup table was created using GHEDesigner to size borehole fields for different thermal loads and ground conditions experienced across the country. A cost and financing structure, along with incentives, were applied. Cost estimates include costs for the distribution network, borehole field installation and operation, and circulation pump operation, while savings are calculated based on energy bills for building owners (agents). This newly developed module can be used for assessing which areas of the country have the highest potential for agent benefits from ambient-temperature loop installation and assess the impact of future cost and price scenarios. Initial results for statewide analysis (for Vermont) and nationwide (for United States) are provided. Future work includes expanding the module to consider mixed residential and commercial districts as well as evaluating multiple cost scenarios.
The Nile basin River system spans 3 million km(exp 2) distributed over ten nations. The eight upstream riparian nations, Ethiopia, Eretria, Uganda, Rwanda, Burundi, Congo, Tanzania and Kenya are the source of approximately 86% of the water inputs to the Nile, while the two downstream riparian countries Sudan and Egypt, presently rely on the river's flow for most of the their needs. Both climate and agriculture contribute to the complicated nature of Nile River management: precipitation in the headwaters regions of Ethiopia and Lake Victoria is variable on a seasonal and inter-annual basis, while demand for irrigation water in the arid downstream region is consistently high. The Nile is, perhaps, one of the most difficult trans-boundary water issue in the world, and this study would be the first initiative to combine NASA satellite observations with the hydrologic models study the overall water balance in a to comprehensive manner. The cornerstone application of NASA's Earth Science Research Results under this project are the NASA Land Data Assimilation System (LDAS) and the USDA Atmosphere-land Exchange Inverse (ALEXI) model. These two complementary research results are methodologically independent methods for using NASA observations to support water resource analysis in data poor regions. Where an LDAS uses multiple sources of satellite data to inform prognostic simulations of hydrological process, ALEXI diagnoses evapotranspiration and water stress on the basis of thermal infrared satellite imagery. Specifically, this work integrates NASA Land Data Assimilation systems into the water management decision support systems that member countries of the Nile Basin Initiative (NBI) and Regional Center for Mapping of Resources for Development (RCMRD, located in Nairobi, Kenya) use in water resource analysis, agricultural planning, and acute drought response to support sustainable development of Nile Basin water resources. The project is motivated by the recognition that accurate, frequent, and spatially distributed estimates of the water balance are necessary for effective water management. This creates a challenge for watersheds that are large, include data poor regions, and/or span multiple nations. All of these descriptors apply to the Nile River basin, yet successful management of the Nile is critical for development and political stability in the region. For this reason, improved hydrological data to support cooperative water management in the Nile basin is a priority for USAID, the US State Department, the World Bank and other international organizations. In this project, the U.S. based research team is working with partners at RCMRD, Nile Basin Initiative (NBI), and their member national-level agencies to develop satellite-based land cover maps, satellite-derived evapotranspiration estimates (using the ALEXI algorithm), and NASA's Land Data Assimilation System (LDAS) customized to match identified information needs. The cornerstone applied sciences product of the project is the development of a customized "Nile LDAS" that will produce optimal estimates of hydrological states and fluxes, as vetted against the in situ observations of NBI and RCMRD member organizations and independent satellite-derived hydrological estimates. Nile LDAS will be applied to improve the reliability of emerging Decision Support Systems in applications that include drought monitoring, reservoir management, and irrigation planning. The end-users such as RCMRD, NBI, Ethiopian and Kenya Meteorological and Famine Early Warning System Network (FEWSNet) will be the eventual benefactors of this work. There will be a capacity building process involving the above end-user organizations and transfer the models and the results for these organizations to execute for future use. The team has already initiated this study and the early results of first years' work are shown. The plan is to complete this work by late 2013.
This paper describes an approach to ontology negotiation between information agents. Ontologies are declarative (data driven) expressions of an agent's "world": the objects, operations, facts, and rules that constitute the logical space within which an agent performs. Ontology negotiation enables agents to cooperate in performing a task, even if they are based on different ontologies. 'Me process allows agents to discover ontology conflicts and then, though incremental interpretation, clarification, and explanation, establish a common basis for communicating with each other. The need for ontology negotiation stems from the proliferation of information sources and of agents with widely varying specialty expertise. The unmanageability of massive amounts of web-based information is already becoming apparent. It is starting to have an impact on professions that rely on distributed archived information. If the expansion continues at its present rate without an ontology negotiation process being introduced, there will soon be no way to ensure the accuracy and completeness of information that scientists obtain from sources other than their own experiments. Ontology negotiation is becoming increasingly recognized as a crucial element of scalable agent technology. This is because agents, by their very nature, are supposed to operate with a fair amount of autonomy and independence from their end-users. Part of this independence is the ability to enlist other agents for help in performing a task (such as locating information on the web). The agents enlisted for help may be "owned" by a different end-user or organization (such as a document archive), and there is no guarantee that they will use the same terminology or understand the same concepts (objects, operators, theorems, rules) as the recruiting agent. For NASA, the need for ontology negotiation arises at the boundaries between scientific disciplines. For example: modeling the effects of global warming might involve knowledge about imaging, climate analysis, ecology, demographics, industrial economics, and biology. The need for ontology negotiation also arises at the boundaries between scientific programs. For example, a Principal Investigator may want to use information from a previous mission to complement downloads from the instruments currently deployed.
The cost-effective and sustainable deployment of hydrogen supply and demand networks, especially in large economic regions like California, can be challenging considering the spatial-temporal availability and variability of the different actors across the network such as production processes, distribution modes, and end-users. In this presentation, we will provide an overview and demonstration of a modeling framework used to assess the environmental, economic, and human health impacts of plausible hydrogen production and supply chain networks in California. Scenarios focus on green hydrogen production pathways using water electrolysis and biomass gasification. End-use applications included in the model are transit, medium and heavy-duty trucking, port authorities, and power and aviation companies that currently consume natural gas, diesel, and aviation fuel for their day-to-day operation. Representative locations for hydrogen production and end-use are based on recent projections of the hydrogen economy in California. All mass and energy flows, as well as estimated emissions, are based on H2A process model designs and projections of technology performance, literature review, and LBNL process, economic and life cycle modeling, and not on company data for the sake of this presentation. Human health impacts are included following methodologies developed for the University of California Irvine HyDeal project. Life cycle phases associated with hydrogen production include feedstock preparation (water and biomass), energy production and consumption (renewable, grid, and combination of renewable and grid electricity), maintenance (chemical utilization in electrolysis and natural gas combustion in gasification), carbon sequestration, hydrogen storage (compression and liquefaction), and distribution (truck and pipeline). We apply the framework utilizing California specific emission factors, financial data, and human health damages and explore the impact of network characteristics on results. Example variations include: the inclusion of policy incentives or not, different representations of the electricity grid and source, electrolysis versus gasification versus combinations of both for production, liquefaction versus compression based on producer capacity cutoffs, transportation truck versus pipeline based on existing infrastructure, and ultimate end use. Comparison of these different scenarios can help inform future projects by demonstrating the trade-offs among environmental, economic, and human health impacts. This model, automated in R, is a starting platform upon which new analysis, modeling capabilities, locations, and emission factors can be rapidly tested and integrated.
A user requirements analysis (URA) was undertaken to determine and appropriate public domain Geographic Information System (GIS) software package for potential integration with NASA's LAS (Land Analysis System) 5.0 image processing system. The necessity for a public domain system was underscored due to the perceived need for source code access and flexibility in tailoring the GIS system to the needs of a heterogenous group of end-users, and to specific constraints imposed by LAS and its user interface, Transportable Applications Executive (TAE). Subsequently, a review was conducted of a variety of public domain GIS candidates, including GRASS 3.0, MOSS, IEMIS, and two university-based packages, IDRISI and KBGIS. The review method was a modified version of the GIS evaluation process, development by the Federal Interagency Coordinating Committee on Digital Cartography. One IEMIS-derivative product, the ALBE (AirLand Battlefield Environment) GIS, emerged as the most promising candidate for integration with LAS. IEMIS (Integrated Emergency Management Information System) was developed by the Federal Emergency Management Agency (FEMA). ALBE GIS is currently under development at the Pacific Northwest Laboratory under contract with the U.S. Army Corps of Engineers' Engineering Topographic Laboratory (ETL). Accordingly, recommendations are offered with respect to a potential LAS/ALBE GIS linkage and with respect to further system enhancements, including coordination with the development of the Spatial Analysis and Modeling System (SAMS) GIS in Goddard's IDM (Intelligent Data Management) developments in Goddard's National Space Science Data Center.