Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “information management and knowledge discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Interactive Web Application for Traffic Simulation Data Management and Visualization

As traffic simulation software becomes more effective for realistically simulating and analyzing traffic dynamics and vehicle interactions on the mesoscopic and microscopic level, the management, dissemination, and collaborative visualization of traffic simulation results produced by individual transportation planners presents a significant challenge. Existing online content management systems have a very limited capability in allowing users to query specific traffic simulation scenarios and geospatially visualize simulation results through shareable and interactive web interfaces. This paper presents a web-based application for promoting the archiving, sharing, and visualization of large-scale traffic simulation outputs. The application is developed to enhance cyber-physical controls, communications, and public education for collaborative transportation planning. Unique features of the web application include: (a) allowing users to upload their new traffic simulation scenarios (parameters and outputs), as well as search existing scenarios using easily accessible interfaces; (b) optimizing simulation output files with heterogeneous data formats and projected coordinate systems for web-based storage and management using a scalable and searchable data/metadata standard; (c) standardizing user-uploaded simulation outputs using web interfaces and data processing libraries with parallel computing capacity; and (d) providing shareable web visual interfaces for visualizing the traffic flow and signal information stored in simulation outputs (e.g., regional traffic patterns and individual vehicle interactions) and visually comparing multiple simulation outputs both spatially and temporally. Furthermore, the paper presents the conceptual design and implementation of this application, and demonstrates the application’s performance for sharing, comparing, and visualizing simulation outputs from VISSIM and SUMO, two commonly used traffic simulation software programs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM). Final Report

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM) is a proposed system that will enable data from an infinite number of diverse sources to be organized and accessed from anywhere using any handheld or other computer device. The approach offers a powerful roadmap for the creation and integration of a unified knowledge base of an entire ecosystem, including its many geophysical, geographical, social, political, agricultural, energy, transportation, and cyber aspects. The resulting aggregation of data has the potential to generate an informational universe of unprecedented size that has never before been possible due to the prohibitive costs, managerial complexity, and technical barriers associated with ever-changing exponential-growth data flows. We envision that DREAM will accelerate discovery by enabling climate researchers, among other types of researchers, to manage, analyze, and visualize data from earth-scale measurements and simulations. DREAM’s success will be built on proven components that leverage existing services and resources. A key building block for DREAM will be the ESGF, chaired by Dean N. Williams. Expanding on the existing ESGF, the project will ensure that the access, storage, movement, and analysis of the large quantities of data that are processed and produced by diverse science projects can be dynamically distributed with proper resource management. Much of the Office of Science data is currently generated by multiple stand-alone facilities. DREAM can collect data accumulated from these facilities and incorporate it into a fully integrated network accessible from anywhere in the world. The result is a completely new paradigm shift for data management, analysis, and visualization enabling researchers to: Manage their calculations, data, tools, and research results; Ensure that all data are sharable, reproducible and (re)usable—accompanied by appropriate metadata describing its provenance, syntax, and semantics at creation; Advance application performance by selectively adapting APIs and services in response to scientific requirements and architectural complexities; and Provide scalable interactive resource management—navigate data and metadata at multiple levels, provide architecture-aware data integration, analysis and visualization tools. We will engage closely with DOE, NASA, and NOAA science groups working at the leading edge of computing. These engagements—in domains such as biology, climate, and hydrology—will allow us to advance disciplinary science goals and inform our development of technologies that can accelerate discovery across DOE more broadly. We will advertise and promote our technologies via dedicated workshops, tutorials, and sessions at conferences, stand-alone events with broad inter-disciplinary invitation, and engagements with leadership facilities.

54 ENVIRONMENTAL SCIENCES↗

Ontologies: The Gateway to Knowledge-Enabled Information Services at Los Alamos National Laboratory

At LANL, we use ontologies to capture and maintain essential organizational knowledge and support tools and frameworks for information discovery. Ontologies capture knowledge, creating meaningful structures for finding and interpreting information. Information in a human context that provides meaning and data that is organized and communicated. This report provides information about LANL's ontology efforts and challenges.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Aligning Standards Communities for Omics Biodiversity Data: Sustainable Darwin Core-MIxS Interoperability

The standardization of data, encompassing both primary and contextual information (metadata), plays a pivotal role in facilitating data (re-)use, integration, and knowledge generation. However, the biodiversity and omics communities, converging on omics biodiversity data, have historically developed and adopted their own distinct standards, hindering effective (meta)data integration and collaboration. In response to this challenge, the Task Group (TG) for Sustainable DwC-MIxS Interoperability was established. Convening experts from the Biodiversity Information Standards (TDWG) and the Genomic Standards Consortium (GSC) alongside external stakeholders, the TG aimed to promote sustainable interoperability between the Minimum Information about any (x) Sequence (MIxS) and Darwin Core (DwC) specifications. To achieve this goal, the TG utilized the Simple Standard for Sharing Ontology Mappings (SSSOM) to create a comprehensive mapping of DwC keys to MIxS keys. This mapping, combined with the development of the MIxS-DwC extension, enables the incorporation of MIxS core terms into DwC-compliant metadata records, facilitating seamless data exchange between MIxS and DwC user communities. Through the implementation of this translation layer, data produced in either MIxS- or DwC-compliant formats can now be efficiently brokered, breaking down silos and fostering closer collaboration between the biodiversity and omics communities. To ensure its sustainability and lasting impact, TDWG and GSC have both signed a Memorandum of Understanding (MoU) on creating a continuous model to synchronize their standards. These achievements mark a significant step forward in enhancing data sharing and utilization across domains, thereby unlocking new opportunities for scientific discovery and advancement.

59 BASIC BIOLOGICAL SCIENCES↗

Development of a Framework for Data Integration, Assimilation, and Learning for Geological Carbon Sequestration (DIAL-GCS) (Final Report)

This project aimed to develop and demonstrate a Data Integration, Assimilation, and Learning framework for geologic carbon sequestration projects (DIAL-GCS). DIAL-GCS is an intelligence monitoring system (IMS) for automating GCS closed-loop management by leveraging recent developments in machine learning technologies, complex event processing (CEP), and reduced-order modeling. The safe and efficient operation of GCS repositories requires integrated monitoring to track the injected CO¬2 as it moves within a storage reservoir. GCS projects are data intensive, as a result of proliferation of digital instrumentation and smart-sensing technologies. GCS projects are also resource intensive, often requiring multidisciplinary teams performing different monitoring, verification, accounting (MVA) tasks throughout the lifecycle of a project to ensure secure containment of injected CO2. The success of GCS thus depends in a large part on our ability to access, assimilate, and analyze heterogeneous data and information sources in a timely manner. This project included a number of meaningful and necessary tasks to transform the human domain knowledge into machine-interpretable rules for automating knowledge extraction and discovery in GCS. The specific technical objectives of the proposed DIAL-GCS project were to develop an ontology-driven GCS data management module for storing, querying, and exchanging GCS data (both historic and live sensor data) from multiple sources and in heterogeneous formats. Incorporate a CEP engine for detecting abnormal situations by seamlessly combining expert knowledge, rule-based reasoning, and machine learning. Enable uncertainty quantification and predictive analytics using a combination of coupled-process modeling, AI/ML methods, and reduced-order modeling, and integrate and demonstrate the system’s capabilities with both real and simulated data. As far as we know, this is one of the first projects aimed to develop intelligent monitoring systems (IMS) targeting the GCS. Under this project, the team had developed a large number of web applications and scientific algorithms that contribute the main theme of intelligent monitoring. The team has published more than a dozen peer reviewed papers and disseminated the research results at multiple technical meetings.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Radiological Recovery Logistics Tool - 20161

Argonne is building and testing a tool, the Radiological Recovery Logistics Tool (RRLT), that can be used during the response and recovery from a radiological or nuclear incident to effectively allocate appropriate commercial and public works equipment to mitigate, remove, and contain radiological contamination. The requirements for this tool - as well as development of the resulting software - is overseen by a steering committee of stakeholders from DHS's National Urban Security Technology Laboratory (NUSTL), the Federal Emergency Management Agency (FEMA), and the Environmental Protection Agency (EPA). One essential requirement is for RRLT to support the efficient and appropriate allocation of resources for a radiological response. Subsequent discussions between ANL and stakeholders have solidified the nature of this support to include identification of the types of resources to be allocated. The study reported in this paper has both factored fundamental concepts and connections out of this identification process and created a Knowledge Base detailing support goals, response scenarios, and efficacy information on dozens of equipment types. In short, RRLT will dynamically apply these findings to situational conditions surrounding contamination incidents. RRLT's Domain, the model of elements, ideas and relationships with which the tool will work, draws concepts from technical reports and stakeholder vocabularies to connect response goals and scenarios to types of equipment that offer utility towards those goals in those scenarios. RRLT's Knowledge Base will contain details on dozens of equipment types and facilitate the operator's discovery and consumption of these details most pertinent to a dynamically selected subset of goals. The core of its Domain Model is based on a report authored by this team. This report [1] contains a comprehensive list of proposed equipment to accomplish various missions or scenarios that might arise after a large-scale radiological contamination incident in an urban environment or critical infrastructure. The report divides potential response and recovery efforts into five support goals: Survey and monitoring of the contaminated area; Mitigation of received dose to first responders: Decontamination (gross and final) of buildings, vehicles, roadways, parks, and other surfaces: Waste management of solid waste generated during recovery operations: and Containment of wastewater and other waste generated during the response and recovery phases. RRLT's development is driven by use cases. A use case is an intention with which a user approaches the software. Use cases are grouped into delivery increments to schedule development, testing, and presentation to stakeholders. This model partitions the system into seven increments: User Arrival and Authentication, Search and Navigation, Equipment Recommendation, Plan Management, Content Management, and Expanded Access. Once a user 15 authenticated, RRLT will present the user with a dashboard that allows them to explore or search RRLT's content. The dashboard will also include a 'Plan' panel for collecting decisions and relevant observations about an incident at hand to facilitate development of an equipment list. RRLT will offer three general modes of access to items in the knowledge base: - Keyword search for direct discovery of items, - Navigation along predetermined paths from recovery goal towards equipment types, and - Interactive guidance towards equipment types by an autonomous software agent: the Equipment Recommendation Wizard. This presentation will detail progress in the development of the RRLT and also discuss opportunities for those interested in providing feedback on its content and functionality. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Knowledge Beacons: Web services for data harvesting of distributed biomedical knowledge

The continually expanding distributed global compendium of biomedical knowledge is diffuse, heterogeneous and huge, posing a serious challenge for biomedical researchers in knowledge harvesting: accessing, compiling, integrating and interpreting data, information and knowledge. In order to accelerate research towards effective medical treatments and optimizing health, it is critical that efficient and automated tools for identifying key research concepts and their experimentally discovered interrelationships are developed. As an activity within the feasibility phase of a project called “Translator” (https://ncats.nih.gov/translator) funded by the National Center for Advancing Translational Sciences (NCATS) to develop a biomedical science knowledge management platform, we designed a Representational State Transfer (REST) web services Application Programming Interface (API) specification, which we call a Knowledge Beacon. Knowledge Beacons provide a standardized basic API for the discovery of concepts, their relationships and associated supporting evidence from distributed online repositories of biomedical knowledge. This specification also enforces the annotation of knowledge concepts and statements to the NCATS endorsed the Biolink Model data model and semantic encoding standards (https://biolink.github.io/biolink-model/). Implementation of this API on top of diverse knowledge sources potentially enables their uniform integration behind client software which will facilitate research access and integration of biomedical knowledge.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Snowmass Computational Frontier: Topical Group Report on Quantum Computing

Quantum computing will play a pivotal role in the High Energy Physics (HEP) science program over the early parts of the 21$^{st}$ Century, both as a major expansion of our capabilities across the Computational Frontier, and in synthesis with quantum sensing and quantum networks. This report outlines how Quantum Information Science (QIS) and HEP are deeply intertwined endeavors that benefit enormously from a strong engagement together. Quantum computers do not represent a detour for HEP, rather they are set to become an integral part of our discovery toolkit. Problems ranging from simulating quantum field theories, to fully leveraging the most sensitive sensor suites for new particle searches, and even data analysis will run into limiting bottlenecks if constrained to our current computing paradigms. Easy access to quantum computers is needed to build a deeper understanding of these opportunities. In turn, HEP brings crucial expertise to the national quantum ecosystem in quantum domain knowledge, superconducting technology, cryogenic and fast microelectronics, and massive-scale project management. The role of quantum technologies across the entire economy is expected to grow rapidly over the next decade, so it is important to establish the role of HEP in the efforts surrounding QIS. Fully delivering on the promise of quantum technologies in the HEP science program requires robust support. It is important to both invest in the co-design opportunities afforded by the broader quantum computing ecosystem and leverage HEP strengths with the goal of designing quantum computers tailored to HEP science.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

ET-AL: Entropy-targeted active learning for bias mitigation in materials data

Growing materials data and data-driven informatics drastically promote the discovery and design of materials. While there are significant advancements in data-driven models, the quality of data resources is less studied despite its huge impact on model performance. In this work, we focus on data bias arising from uneven coverage of materials families in existing knowledge. Observing different diversities among crystal systems in common materials databases, we propose an information entropy-based metric for measuring this bias. To mitigate the bias, we develop an entropy-targeted active learning (ET-AL) framework, which guides the acquisition of new data to improve the diversity of underrepresented crystal systems. We demonstrate the capability of ET-AL for bias mitigation and the resulting improvement in downstream machine learning models. This approach is broadly applicable to data-driven materials discovery, including autonomous data acquisition and dataset trimming to reduce bias, as well as data-driven informatics in other scientific domains.

36 MATERIALS SCIENCE↗

DNA Sequence-Based Identification of Fusarium : A Work in Progress

Accurate species-level identification of an etiological agent is crucial for disease diagnosis and management because knowing the agent’s identity connects it with what is known about its host range, geographic distribution, and toxin production potential. This is particularly true in publishing peer-reviewed disease reports, where imprecise and/or incorrect identifications weaken the public knowledge base. This can be a daunting task for phytopathologists and other applied biologists that need to identify Fusarium in particular, because published and ongoing multilocus molecular systematic studies have highlighted several confounding issues. Paramount among these are: (i) this agriculturally and clinically important genus is currently estimated to comprise more than 400 phylogenetically distinct species (i.e., phylospecies), with more than 80% of these discovered within the past 25 years; (ii) approximately one-third of the phylospecies have not been formally described; (iii) morphology alone is inadequate to distinguish most of these species from one another; and (iv) the current rapid discovery of novel fusaria from pathogen surveys and accompanying impact on the taxonomic landscape is expected to continue well into the foreseeable future. To address the critical need for accurate pathogen identification, our research groups are focused on populating two web-accessible databases (FUSARIUM-ID v.3.0 and the nonredundant National Center for Biotechnology Information nucleotide collection that includes GenBank) with portions of three phylogenetically informative genes (i.e., TEF1, RPB1, and RPB2) that resolve at or near the species level in every Fusarium species. The objectives of this Special Report, and its companion in this issue ( Torres-Cruz et al. 2022 ), are to provide a progress report on our efforts to populate these databases and to outline a set of best practices for DNA sequence-based identification of fusaria.

Plant Sciences↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

Plant science decadal vision 2020–2030: Reimagining the potential of plants for a healthy and sustainable future

Abstract Plants, and the biological systems around them, are key to the future health of the planet and its inhabitants. The Plant Science Decadal Vision 2020–2030 frames our ability to perform vital and far‐reaching research in plant systems sciences, essential to how we value participants and apply emerging technologies. We outline a comprehensive vision for addressing some of our most pressing global problems through discovery, practical applications, and education. The Decadal Vision was developed by the participants at the Plant Summit 2019, a community event organized by the Plant Science Research Network. The Decadal Vision describes a holistic vision for the next decade of plant science that blends recommendations for research, people, and technology. Going beyond discoveries and applications, we, the plant science community, must implement bold, innovative changes to research cultures and training paradigms in this era of automation, virtualization, and the looming shadow of climate change. Our vision and hopes for the next decade are encapsulated in the phrase reimagining the potential of plants for a healthy and sustainable future. The Decadal Vision recognizes the vital intersection of human and scientific elements and demands an integrated implementation of strategies for research (Goals 1–4), people (Goals 5 and 6), and technology (Goals 7 and 8). This report is intended to help inspire and guide the research community, scientific societies, federal funding agencies, private philanthropies, corporations, educators, entrepreneurs, and early career researchers over the next 10 years. The research encompass experimental and computational approaches to understanding and predicting ecosystem behavior; novel production systems for food, feed, and fiber with greater crop diversity, efficiency, productivity, and resilience that improve ecosystem health; approaches to realize the potential for advances in nutrition, discovery and engineering of plant‐based medicines, and "green infrastructure." Launching the Transparent Plant will use experimental and computational approaches to break down the phytobiome into a "parts store" that supports tinkering and supports query, prediction, and rapid‐response problem solving. Equity, diversity, and inclusion are indispensable cornerstones of realizing our vision. We make recommendations around funding and systems that support customized professional development. Plant systems are frequently taken for granted therefore we make recommendations to improve plant awareness and community science programs to increase understanding of scientific research. We prioritize emerging technologies, focusing on non‐invasive imaging, sensors, and plug‐and‐play portable lab technologies, coupled with enabling computational advances. Plant systems science will benefit from data management and future advances in automation, machine learning, natural language processing, and artificial intelligence‐assisted data integration, pattern identification, and decision making. Implementation of this vision will transform plant systems science and ripple outwards through society and across the globe. Beyond deepening our biological understanding, we envision entirely new applications. We further anticipate a wave of diversification of plant systems practitioners while stimulating community engagement, underpinning increasing entrepreneurship. This surge of engagement and knowledge will help satisfy and stoke people's natural curiosity about the future, and their desire to prepare for it, as they seek fuller information about food, health, climate and ecological systems.

59 BASIC BIOLOGICAL SCIENCES↗

South Dakota Region Scientific Deep Dive

EPOC uses the Deep Dive process to discuss and analyze current and planned science use cases and anticipated data output of a particular use case, site, or project to help inform the strategic planning of a campus or regional networking environment. This includes understanding future needs related to network operations, network capacity upgrades, and other technological service investments. A Deep Dive comprehensively surveys major research stakeholders’ plans and processes in order to investigate data management requirements over the next 5–10 years. Questions crafted to explore this space include the following: 1) How, and where, will new data be analyzed and used? 2) How will the process of doing science change over the next 5–10 years? and 3) How will changes to the underlying hardware and software technologies influence scientific discovery? Deep Dives help ensure that key stakeholders have a common understanding of the issues and the actions that a campus or regional network may need to undertake to offer solutions. The EPOC team leads the effort and relies on collaboration with the hosting site or network, and other affiliated entities that participate in the process. EPOC organizes, convenes, executes, and shares the outcomes of the review with all stakeholders

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Dynamic and Responsive Distributed Energy Resource Education Solutions for Building, Fire, and Safety Department Officials (Final Technical Report)

From April 2021 through March 2024, the Interstate Renewable Energy Council (IREC) led a collaborative project to develop a free online clearinghouse of educational resources about solar photovoltaics (PV), energy storage systems (ESS), electric vehicle supply equipment (EVSE), and grid-interactive efficient building (GEB) technologies. Two websites—the Clean Energy Clearinghouse and CleanEnergyTraining.org—housed over 70 educational resources. Over the course of the three-year project, 154,272 unique visitors accessed the learning materials. Learner feedback was overwhelmingly positive. Even through the end of the project, there was sustained demand for education and communication. A primary innovation of the project was to drive multiple complementary audiences to the same place. Building owners, designers, installation contractors and developers, authorities having jurisdiction (AHJs), and fire service personnel all benefit from a shared understanding of clean energy technologies, including safety and code-related requirements. When considering the impact on the target audience, the project team worked with partners and advisors to inform resource creation and delivery in such a way as to address key motivational factors of the target audience and compel each user to seek additional information on the topic and return to the Clean Energy Clearinghouse website as their central location for more information. Resources were intentionally developed to be concise—five to 15 minutes—and accessible, meaning not overly technical. Providing basic information demystified the technologies and invited the professional to explore additional learning opportunities. Awardee and partner collaboration was key to project success. IREC facilitated collaboration among the other Topic 2 awardees, Southface and New Buildings Institute (NBI). The three awardees shared relevant information gained through discovery and validation questionnaires that informed product development and reduced duplication of effort by coordinating the development of complementary, and not competing, educational resources. Inspired by this collaboration, IREC brought on additional partners even in the final year of the project. Five regional energy efficiency organizations were part of the project, which expanded the connection between efficiency and distributed energy resources. We also included resources on the Clearinghouse that were developed through other federally funded projects, such as the Buildings Energy Efficiency Frontiers & Innovation Technologies (BENEFIT) program. The website was developed with the learner in mind, and not solely the funding source. Feedback from stakeholders throughout the project, and especially in its final year, indicated the need for continued education and facilitated communication among stakeholders to further the safe and widespread adoption of clean energy.

14 SOLAR ENERGY↗

Challenges and Advances in Information Extraction from Scientific Literature: a Review

Scientific articles have long been the primary means of disseminating scientific discoveries. Over the centuries, valuable data and potentially groundbreaking insights have been collected and buried deep in the mountain of publications. In materials engineering, such data are spread across technical handbooks specification sheets, journal articles, and laboratory notebooks in myriad formats. Extracting information from papers on a large scale has been a tedious and time-consuming job to which few researchers have wanted to devote their limited time and effort, yet is an activity that is essential for modern data-driven design practices. However, in recent years, significant progress has been made by the computer science community on techniques for automated information extraction from free text. Yet, transformative application of these techniques to scientific literature remains elusive-due not to a lack of interest or effort but to technical and logistical challenges. Using the challenges in the materials science literature as a driving motivation, we review the gaps between state-of-the-art information extraction methods and the practical application of such methods to scientific texts, and offer a comprehensive overview of work that can be undertaken to close these gaps.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗