Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “natural user interfaces”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ReVise: A Human-AI Interface for Incremental Algorithmic Recourse

The recent adoption of artificial intelligence in socio-technical systems raises concerns about the black-box nature of the resulting decisions in fields such as hiring, finance, admissions, etc. If data subjects—such as job applicants, loan applicants, and students—receive an unfavorable outcome, they may be interested in algorithmic recourse, which involves updating certain features to yield a more favorable result when re-evaluated by algorithmic decision-making. Unfortunately, when individuals do not fully understand the incremental steps needed to change their circumstances, they risk following misguided paths that can lead to significant, long-term adverse consequences. Existing recourse approaches focus exclusively on the final recourse goal but neglect the possible incremental steps to reach the goal with real-life constraints, user preferences, and model artifacts. To address this gap, we formulate a visual analytic workflow for incremental recourse planning in collaboration with AI/ML experts and contribute an interactive visualization interface that helps data subjects efficiently navigate the recourse alternatives and make an informed decision. We also present one of the many usage scenarios, developed during exploratory feedback sessions with twelve graduate students using a real-world dataset, which demonstrates that our approach can be instrumental for data subjects in choosing a suitable recourse path.

algorithmic recourse↗

AFUE Analysis Software Tool

The Annual Fuel Utilization Efficiency (AFUE) analysis of residential and light commercial furnaces follows ANSI/ASHRAE Standard 103-2017 (i.e., Method of Testing for Annual Fuel Utilization Efficiency of Residential Central Furnaces and Boilers) . The analysis is a complex, comprehensive method based on furnace configuration and specific components equipped, requiring detailed furnace testing and measurement data. For this reason, an AFUE Analysis Tool using Microsoft Excel enabled with Visual Basic for Applications (VBA) was developed with a user-friendly interface and comprehensive coverage. The tool consists of three worksheets: unit and configuration selection, geometry and measurement data input, and AFUE plus key results. This tool can be used to estimate the AFUE of both condensing and noncondensing furnaces with single-stage, two-stage, and step-modulating functions. The tool was validated with experimental data from Oak Ridge National Laboratory’s natural gas furnace projects that are commercially available. The results indicate the tool is reasonably accurate in the evaluation of a new R&D modified furnace unit.

Gao, Zhiming [Oak Ridge National Laboratory (ORNL)↗

Multi-Scale Integrated Monitoring System for Enhancing Methane Emission Detection, Quantification & Prediction

This report details the progress and findings of a comprehensive study on reviewing existing solutions, identifying technology gaps, and formulating an “all-in-one” integrated strategy for developing the next-generation multiscale methane monitoring and modeling platform, conducted under grant number DE-FE0032292. Co-led by Dr. David Ebert, Dr. Binbin Weng, and Dr. Chenghao Wang at the University of Oklahoma, the project’s goal was to develop an integrated approach for building this engineering platform to detect, quantify, and mitigate methane emissions across various temporal scale, spatial scales, and sectors. The planning grant study began with an extensive review of various methane sensing and monitoring technologies and systems, surveying over 100 technology providers globally. This review revealed the prevalence of optical methods over chemical methods in commercially available sensors, with Non-Dispersive Infrared (NDIR), Tunable Diode Laser Absorption Spectroscopy (TDLAS), and Optical Gas Imaging (OGI) cameras being the most prevalent options. A trend towards more advanced optical techniques was observed, driven by increased regulatory focus and technological advancements. The technical evaluation of these sensing technologies provided crucial insights into their capabilities and limitations. The study examined emerging technologies such as Differential Absorption LiDAR (DIAL), which show promise for high-precision and long-range detection. The team then investigated the features and application bandwidth of various sensing platforms, including handheld, fixed/stationary, mobile, aerials, and spaceborne monitors. Pilot field studies were conducted to assess the capabilities of solutions for different emission scenarios. Field work with sensor deployments was conducted at three distinct site types: an oil & gas industry site, a cattle ranching operation, and a waste processing facility. The team also conducted a thorough review of methane flux inverse modeling approaches, focused on physically based methods. These approaches were categorized into simple, intermediate, and advanced methods. A realtime WRF-GHG (Weather Research and Forecasting-Greenhouse Gas) modeling system was developed and applied, incorporating multiple data sources to guide field experiments and inform methane plume detection. The project identified and analyzed numerous categories of methane data sources, including satellite measurements, ground-based sensors, and inventory databases. Key platforms examined include EDGAR, EPA GHGI, NASA TROPOMI, Carbon Mapper, and Climate TRACE, among others. The team proposed an architecture for a comprehensive methane monitoring platform. This system incorporates multi-source data acquisition, advanced data processing and assimilation, interactive visualization tools, and analytical capabilities for emissions forecasting and scenario analysis. The proposed platform aims to provide a user-friendly interface catering to various stakeholders, from researchers to policymakers. The architecture includes sophisticated data ingestion methods, a centralized data warehouse, and advanced analytical tools for data fusion and interpretation. To ensure the relevance and effectiveness of the proposed system, a comprehensive survey was conducted to gather stakeholder input on system requirements. Key findings include a strong need for integrating various data types and formats, a preference for real-time data updates and advanced visualization tools, and a demand for user-friendly interfaces catering to different expertise levels.

03 NATURAL GAS↗

Aquifer Injection Modeling (AIM) Toolbox User Guide

The Aquifer Injection Modeling Toolbox (“AIM Toolbox”) software was developed by the Pacific Northwest National Laboratory (PNNL) for the U.S. Environmental Protection Agency’s (EPA) to provide a collection of analytical solutions suitable for evaluating the potential extent of the area impacted by subsurface injection operations. Subsurface injection operations are regulated under the EPA’s Underground Injection Control (UIC) program and are typically related to oil/gas development, waste disposal, or subsurface mining or storage. The analytical algorithms provided in the AIM Toolbox each have different approaches/assumptions/focus with respect to the nature and processes in subsurface and the nature of the injection operations. Collectively the set of analysis algorithms provides a broader evaluation of the area that can potentially be impacted by an injection operation. By providing estimates for the injectate plume extent, the software supports technical aspects of planning, evaluation, and overseeing injection activities. That is, the results help assess when further regulatory controls (e.g., monitoring, reporting) may be required and, in the case of disposal into an underground source of drinking water, the extent of the impacted area would require exemption from protection under the Safe Drinking Water Act. The AIM Toolbox software is available as a single-page web application, providing an interface to provide the necessary inputs, and both chart and map panes for visualizing the results. This User Guide describes the AIM Toolbox software, information required, user interactions, and the underlying basis of the calculations.

58 GEOSCIENCES↗

OpenStudio®-MCP [SWR-26-035]

OpenStudio®-MCP is a Model Context Protocol (MCP) server that lets AI assistants perform building energy modeling through natural language. Rather than requiring users to learn the OpenStudio® SDK, EnergyPlus® scripting, or Ruby/Python automation, the server translates conversational requests into sequences of tool calls that create models, design HVAC systems, run simulations, and extract results — all within a single chat session. The server's 124 tools are organized into a skills architecture where each skill encapsulates a domain of building energy modeling (envelope, HVAC, loads, weather, simulation, results) behind typed, LLM-friendly interfaces. High-leverage operations like applying ASHRAE 90.1 baseline systems or generating standards-compliant typical buildings are exposed as single tool calls that internally wire dozens of OpenStudio® objects. Bundled measures from ComStock™ and Openstudio® -common-measures-gem are wrapped with dedicated tools and typed arguments rather than exposed through a generic measure interface, so AI models get consistent, error-resistant recipes without needing to discover measure arguments at runtime. A key design decision is structured results extraction: six SQL-based tools return surgical ~300–1,000 token responses (end-use breakdowns, envelope summaries, HVAC sizing, timeseries data) instead of requiring the AI to parse ~100K-token raw HTML reports, making iterative design exploration practical within context window limits. The codebase is designed as a reference implementation — explicit, well-commented, and modular — so that other simulation engines (EnergyPlus® standalone, TRNSYS, DOE-2) can use it as a template for building their own MCP servers.

Ball, Brian [National Laboratory of the Rockies (N↗

Developing a Robust Market for CHP: A Plan for Fostering Economic Development, Business Competitiveness and Resiliency in the Mid-Atlantic Region

The Department of Energy’s Mid Atlantic Combined Heat and Power Technical Assistance Partnership (MA CHP TAP) was established to develop public-private partnerships to advance the technology, policies, and programmatic support for combined heat and power (CHP), including its application in microgrids, heat to power and district energy. The MA CHP TAP’s work includes education and outreach as well as technical assistance to a variety of stakeholders including end-users (commercial, industrial, institutional and more), state decision makers, electric and gas utilities, trade associations and non-profit organizations. This assistance includes evaluating the economic, energy, reliability and environmental value of proposed systems. The MA CHP TAP represents the multi-state Mid- Atlantic region and is the CHP expert in the region who provides fact-based, un-biased information on CHP, including technologies, project development, project financing, local electric and natural gas utility interfaces, and related state best practice policies.

20 FOSSIL-FUELED POWER PLANTS↗

Improved Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graphs are ubiquitous in modeling complex systems and representing interactions between entities to uncover structural information of the domain. Traditionally, graph analytics workloads are challenging to efficiently scale (both strong and weak cases) on distributed memory due to the irregular memory-access driven nature (with little or no computations) of the methods. The structure of graphs and their relative distribution over the processing elements poses another level of complexity, making it difficult to attain sustainable scalability across platforms. In this paper, we discuss enhancements to TriC, a distributed-memory implementation of graph triangle counting using Message Passing Interface (MPI), which was featured in the 2020 Graph Challenge competition. We have made some incremental enhancements to TriC, primarily adopting a user-defined buffering strategy to overcome the startup problem for large graphs (by fixing the memory for intermediate data), and experimenting with probabilistic data structures such as bloom filter to improve the query response time for assessing edge existence, at the expense of increasing the overall false positive rate. These adjustments have led to a modest improvements in most cases, as compared to the previous version.

Graph Analytics, HPC↗

Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower (DIVERS-H)

U.S. hydropower plants face potential threats from shrinking water supply, rising demands, and warmer stream temperatures from various causes. Power plant owners, operators, and regulators require new tools to take advantage of and interpret the diverse range of scientific data being produced by both observational methods (for example, satellite, radar, stream gauges) and computer modeling methods that evaluate and predict how earth's dynamic systems (atmosphere, oceans, land surface, and sea ice) are changing and interacting. Combining datasets such as these with AI-based analyses introduces a novel decision support system to help users anticipate and address potential impacts on power generation stations. This new technology has been named DIVERS-H for "Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower." In Phase I, technical feasibility was established with the development and demonstration of all the new technologies that are required. Most notably, DIVERS-H will use new artificial intelligence (AI) methods to capture the complex dynamics of water availability, demand, and environmental changes. In addition, new data management software was developed, and a prototype user interface was implemented as the precursor to a full scale decision support system. With technical research complete, the project focus now shifts to development of a commercial software product to provide users with actionable insight into water availability and the risk/resilience of critical systems at their locations of interest. Although DIVER-H was originally conceived as a tool for hydroelectric power applications, the same underlying technology can be readily applied to other water-consuming systems including coal, natural gas, oil, and nuclear power plants.

Chaudhary, Aashish [Kitware, Inc., Clifton Park, N↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Dual Context: Leveraging Structured Application Context for Code Generation and Runtime Feature Activation via Chat Interfaces

Integrating artificial intelligence (AI) capabilities into software applications typically involves two common paths. For developers, AI assists in generating and documenting source code and other related software engineering efforts. For users, AI assists them through question-and-answer exchanges via chatbots. Both approaches have their value, but neither effectively leverages the modularity of component-based architectures that modern web application frameworks offer. We implement a proof of concept within a centralized suite of applications used for the Atmospheric Radiation Measurement (ARM) Data Center Operational Tools, where we introduce a third integration path through the ARM Context Engine (ACE). ACE is a context driven system that uses structured contextual specifications to enable Large Language Models (LLMs) to render interactive and feature-rich user interface (UI) components directly within chat responses, alongside or in place of conventional text outputs. These specifications serve two important purposes across what we call code context and UI context. Code context provides AI-assisted development tools with structured application knowledge beyond raw code, including component relationships, architectural patterns and schematic information, enabling the generation of consistent, well-structured code. UI context defines the rules for enabling and rendering component features at runtime based on the user's natural language input, allowing end users to activate capabilities such as data export, filtering, and pagination within chat responses, without requiring code changes or redeployment. We demonstrate, through a comparative evaluation against general-purpose AI chatbots, that context-driven component rendering provides interactive capabilities that text-based responses cannot replicate, including deterministic component behavior, application-consistent design language, and on-demand feature activation. A development effort comparison further shows that features that traditionally require multi-step development cycles can be activated with a single naturallanguage request. In this ongoing work, we present ACE as an emerging approach to AI integration that positions modular, well-documented software architecture as the foundation for AI-ready applications. ACE treats context as a shared resource across both development and user-facing AI, bringing cohesion to conventionally disconnected efforts, bridging developer tooling and end-user capabilities within a single framework.

Tadimeti, Vijay [ORNL]↗

ToPolyAgent: AI agents for coarse-grained bead-spring topological polymer simulations

We introduce ToPolyAgent, a multi-agent AI framework for performing coarse-grained molecular dynamics (MD) simulations of topological polymers through natural language instructions. By integrating large language models (LLMs) with domain-specific computational tools, ToPolyAgent supports both interactive and autonomous simulation workflows across diverse polymer architectures, including linear, ring, brush, and star polymers, as well as dendrimers. The system consists of four LLM-powered agents: a Config Agent for generating initial polymer–solvent configurations, a Simulation Agent for executing LAMMPS-based MD simulations and conformational analyses, a Report Agent for compiling markdown reports, and a Workflow Agent for streamlined autonomous operations. Interactive mode incorporates user feedback loops for iterative refinements, while autonomous mode enables end-to-end task execution from detailed prompts. We demonstrate ToPolyAgent's versatility through case studies involving diverse polymer architectures under varying solvent conditions, thermostats, and simulation lengths. Furthermore, we highlight its potential as a research assistant by directing it to investigate the effect of interaction parameters on the linear polymer conformation, and the influence of grafting density on the persistence length of the brush polymer. By coupling natural language interfaces with rigorous simulation tools, ToPolyAgent lowers barriers to complex computational workflows and advances AI-driven materials discovery in polymer science. It lays the foundation for autonomous and extensible multi-agent scientific research ecosystems.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

R “SHINY” GUI DEVELOPMENT FOR URANIUM ISOTOPIC ANALYSIS WITH MATRIX-ASSISTED IONIZATION MASS SPECTROMETRY

The international nuclear safeguards community continues to seek rapid, accurate, and precise characterization capabilities for the in-field measurement of uranium isotopic compositions in nuclear facilities. Mass spectrometry (MS) is considered the “gold standard” for analysis of relatively long-lived actinides such as uranium (U) and plutonium; however, conventional MS analysis often requires time consuming sample preparation and complex analytical methodologies that are difficult to perform in-field or in-facility. Matrix assisted ionization (MAI) is a novel ambient ionization MS technique (i.e., MAI-MS) that potentially addresses these challenges due to the relative simplicity of the ionization phenomenon and ruggedness of ambient MS instrumentation. Savannah River National Laboratory (SRNL, USA) has demonstrated this technique for nanogram-level 235U/238U isotope ratio measurements within seconds, with percent-level analytical uncertainties capable of discriminating depleted, natural, and low-enriched uranium. Current experimental work on developing MAI methods for uranium isotopic analysis has been enabled by parallel development of a comprehensive MAI-MS data analysis suite at SRNL. Development of this bespoke data analysis software was necessary because commercially available ambient MS software is poorly suited for uranium isotope ratio measurement. The effort leverages the power of R, a popular open-source programming language, and Shiny, an R package providing tools for graphical user interface (GUI) and web interface coding. This software allows researchers without any programming experience to harness and utilize R’s considerable data analysis/visualization power.

LaBone, Elizabeth D.↗

Local Weather Station Design and Development for Cost-Effective Environmental Monitoring and Real-Time Data Sharing

Current weather monitoring systems often remain out of reach for small-scale users and local communities due to their high costs and complexity. This paper addresses this significant issue by introducing a cost-effective, easy-to-use local weather station. Utilizing low-cost sensors, this weather station is a pivotal tool in making environmental monitoring more accessible and user-friendly, particularly for those with limited resources. It offers efficient in-site measurements of various environmental parameters, such as temperature, relative humidity, atmospheric pressure, carbon dioxide concentration, and particulate matter, including PM 1, PM 2.5, and PM 10. The findings demonstrate the station’s capability to monitor these variables remotely and provide forecasts with a high degree of accuracy, displaying an error margin of just 0.67%. Furthermore, the station’s use of the Autoregressive Integrated Moving Average (ARIMA) model enables short-term, reliable forecasts crucial for applications in agriculture, transportation, and air quality monitoring. Furthermore, the weather station’s open-source nature significantly enhances environmental monitoring accessibility for smaller users and encourages broader public data sharing. With this approach, crucial in addressing climate change challenges, the station empowers communities to make informed decisions based on real-time data. In designing and developing this low-cost, efficient monitoring system, this work provides a valuable blueprint for future advancements in environmental technologies, emphasizing sustainability. The proposed automatic weather station not only offers an economical solution for environmental monitoring but also features a user-friendly interface for seamless data communication between the sensor platform and end users. This system ensures the transmission of data through various web-based platforms, catering to users with diverse technical backgrounds. Furthermore, by leveraging historical data through the ARIMA model, the station enhances its utility in providing short-term forecasts and supporting critical decision-making processes across different sectors.

54 ENVIRONMENTAL SCIENCES↗

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

A Generalized Transformer-Based Pulse Detection Algorithm

Pulse-like signals are ubiquitous in the field of single molecule analysis, e.g., electrical or optical pulses caused by analyte translocations in nanopores. The primary challenge in processing pulse-like signals is to capture the pulses in noisy backgrounds, but current methods are subjectively based on a user-defined threshold for pulse recognition. Here, we propose a generalized machine-learning based method, named pulse detection transformer (PETR), for pulse detection. PETR determines the start and end time points of individual pulses, thereby singling out pulse segments in a time-sequential trace. It is objective without needing to specify any threshold. It provides a generalized interface for downstream algorithms for specific application scenarios. PETR is validated using both simulated and experimental nanopore translocation data. It returns a competitive performance in detecting pulses through assessing them with several standard metrics. Finally, the generalization nature of the PETR output is demonstrated using two representative algorithms for feature extraction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Computing – Real-Time Data Processing from a Dilution Refrigerator

This poster presents a real-time data visualization system for monitoring and displaying data from a Bluefors control unit connected to a dilution refrigerator. The goal is to provide an aesthetically pleasing and user-friendly web-based interface that continuously updates and displays critical data such as temperature and flow rates at various points within the refrigerator. Leveraging WebSocket technology, the system establishes multiple connections to the dilution refrigerators, enabling simultaneous monitoring of several units. The interactive webpage dynamically updates the data, providing researchers and operators with instant insights into the system's performance. The system's ability to stream and visualize data in real-time enhances the understanding of the dilution refrigerator's behavior, aiding in optimizing its operation and ensuring efficient scientific experiments. The application's web-based nature makes it easily accessible from any device with internet connectivity, promoting seamless collaboration and remote monitoring capabilities. Overall, this poster offers a comprehensive solution for real-time data visualization and analysis of dilution refrigerators, catering to the needs of researchers and scientists working in low-temperature physics and quantum computing fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

BioSiting Tool (BioSiting) v2

The BioSiting Tool provides a geospatial interface for analyzing bioeconomy resources and infrastructure across the continental U.S. The tool integrates empirical and modeled data from a broad range of sources. Bioeconomy resources mapped in the tool include agricultural residues, forest residues, municipal solid waste streams, food waste, manure, fats, oils and greases and potential yields of energy crops. Infrastructure mapped in the tool includes biorefineries, material recovery facilities, anaerobic digesters, wastewater treatment plants, combustion plants, district energy systems, crude oil pipelines, petroleum pipelines, natural gas pipelines, railways and freight terminals. Additional data layers include environmental justice indicators at the census tract level and carbon dioxide geologic storage potential. Users can select a location on the map, define a buffer radius in kilometers and generate an inventory of all bioecomony resources within the buffer zone. Data from the tool can be downloaded from individual buffer zones, or at the state or national level.

Huntington, Tyler↗

PERCEPTIVE: an R shiny $\underline{p}$ipelin$\underline{e}$ for the p$\underline{r}$edi$\underline{c}$tion of $\underline{ep}$igenetic modula$\underline{t}$ors $\underline{i}$n no$\underline{v}$el sp$\underline{e}$cies

Epigenetic processes are central to regulating gene expression, genome stability, and metabolic function across the tree of life; yet, their roles remain underexplored in microalgae, especially as new species continue to be identified and characterized. This is likely due to the cumbersome nature and species-dependent attributes of epigenetic wet-lab methodologies, which preclude the rapid identification of epigenetic modifications and modulators. However, there is high conservation of epigenetic processes from budding yeast to humans; in many cases, one may infer how behavior and function are epigenetically regulated in novel species by identifying epigenetic modulators, or the proteins responsible for conferring epigenetic modifications. Here, to this end, we have developed a graphical software package, titled PERCEPTIVE (pipeline for the prediction of epigenetic modulators in novel species). This platform solely uses the genomic sequence of an algal species, and preexisting information from other model organisms, to predict the epigenetic modulators and associated modifications in algae. Predictions are presented to the user in a graphical interface, which provides literature-based interpretation of results, enabling users to quickly understand potential epigenetic processes in their algal species of interest and plan follow-up experiments. To test PERCEPTIVE, we predicted epigenetic modulators in several feedstock candidate algae species. To validate these predictions, wet-lab studies were performed, including mass spectrometry; these results underscore the high accuracy of PERCEPTIVE predictions. Overall, PERCEPTIVE represents a powerful in silico tool for the research and manipulation of algal species, which does not require a priori knowledge of epigenetics and is accessible to a broad set of investigators.

59 BASIC BIOLOGICAL SCIENCES↗