Engineering PapersSearch

SEARCH · Engineering Papers

Results for “AUTOMATIC DATA PROCESSING SYSTEM”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automated analysis of unlabeled PV data with Solar Data Tools software: Overview and feature updates

Distributed rooftop PV systems: ubiquitous, yet commonly have unlabeled data Difficult or impossible to form a performance index We developed Solar Data Tools (SDT), an open-source Python library for analyzing PV power (and irradiance) time-series data SDT enables analysis of unlabeled PV data—no model, no meteorological data, no performance index required Takes a statistical signal processing approach Data processing steps are largely pre-defined and automatic regardless of system type—from utility tracking systems to multi-pitch rooftop systems

Meyers-Im, Bennet E

Spatiotemporal Automatic Calibration of Infrastructure Lidar, Radar, and Camera with a Global Navigation Satellite System

Robust and accurate perception is important for modern intelligent transportation systems (ITS), which use sensors of various modalities for data fusion to create a digital twin of an intersection. Sensor calibration is an important process that creates a unified coordinate frame for the sensor output data so that it can be used for data fusion. Classical approaches for sensor calibration are time-consuming, require an overlapping field of view for feature matching, and are not feasible for ITS application as they cause disruptions in the flow of traffic. In this paper, we present a spatiotemporal automatic calibration approach to calibrate multiple infrastructure lidar, radar, and cameras installed at a traffic intersection. The approach uses global navigation satellite system (GNSS) positioning information shared by connected vehicles, and when the vehicle is detected by the sensor, we match the sensor detections with the GNSS coordinates. The proposed algorithm is evaluated with a real-world dataset utilizing detections from two radars, cameras, and lidars with a test vehicle instrumented with a post-processing kinematic (PPK)-corrected GNSS driving past the sensors installed at a four-way traffic intersection. The experimental results show that the proposed automatic calibration approach can achieve the transformation with a root mean squared error of less than 0.5 for radar and lidar and less than 2 for camera detections. The ability to rapidly calibrate sensors not only benefits initial installations, but can also be used for system health monitoring, while utilizing available connected vehicle data to test the real-time sensor fidelity and operational status.

ADVANCED PROPULSION SYSTEMS,ENERGY CONSERVATION, C

Spatiotemporal Automatic Calibration of Infrastructure Lidar, Radar, and Camera with a Global Navigation Satellite System: Preprint

Robust and accurate perception is important for modern intelligent transportation systems (ITS), which use sensors of various modalities for data fusion to create a digital twin of an intersection. Sensor calibration is an important process that creates a unified coordinate frame for the sensor output data so that it can be used for data fusion. Classical approaches for sensor calibration are time-consuming, require an overlapping field of view for feature matching, and are not feasible for ITS application as they cause disruptions in the flow of traffic. In this paper, we present a spatiotemporal automatic calibration approach to calibrate multiple infrastructure lidar, radar, and cameras installed at a traffic intersection. The approach uses global navigation satellite system (GNSS) positioning information shared by connected vehicles, and when the vehicle is detected by the sensor, we match the sensor detections with the GNSS coordinates. The proposed algorithm is evaluated with a real-world dataset utilizing detections from two radars, cameras, and lidars with a test vehicle instrumented with a post-processing kinematic (PPK)-corrected GNSS driving past the sensors installed at a four-way traffic intersection. The experimental results show that the proposed automatic calibration approach can achieve the transformation with a root mean squared error of less than 0.5 for radar and lidar and less than 2 for camera detections. The ability to rapidly calibrate sensors not only benefits initial installations, but can also be used for system health monitoring, while utilizing available connected vehicle data to test the real-time sensor fidelity and operational status.

ADVANCED PROPULSION SYSTEMS

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Conference presentation at Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Extended abstract for Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM

ACDC (Automated Campbell Diagram Code) [SWR-26-042]

This application provides a web-based graphical user interface to generating Campbell Diagrams and visualizing mode shapes for OpenFAST turbine models. Determining the aeroelastic stability and dynamic characteristics of wind turbines is a critical step in turbine design and analysis. Historically, extracting natural frequencies and mode shapes from OpenFAST—the industry-standard whole-turbine simulation code—has been a fragmented and tedious process. It required manual model configuration, command-line linearization execution, and complex post-processing via proprietary scripts to handle rotating-frame dynamics. To address these workflow bottlenecks, we present the Automated Campbell Diagram Code (ACDC), an open-source graphical software tool developed by the National Laboratory of the Rockies (NLR) under the DOE-funded Distributed Wind Aeroelastic Modeling (dWAM) project. ACDC streamlines the end-to-end linearization and stability analysis workflow into a single, intuitive cross-platform application. The software guides users through OpenFAST model configuration, definition of operating points, and the automated execution of steady-state trim and linearization simulations. Under the hood, ACDC automates the complex mathematical post-processing steps required for rotating systems, including Multi-Blade Coordinate (MBC) transformations, eigenanalysis, and advanced modal tracking utilizing the Modal Assurance Criterion (MAC) and spectral clustering. Finally, ACDC processes these results to automatically generate Campbell diagrams and features a robust 3D visualization engine to animate full-system mode shapes. By eliminating the reliance on external post-processing environments and manual data manipulation, ACDC significantly accelerates dynamic analysis and lowers the barrier to entry for wind energy researchers and engineers.

Summerville, Brent [National Laboratory of the Roc

SolarAPP+ Performance Review (2024 Data)

The Solar Automated Permit Processing Plus (SolarAPP+) platform is an online portal to facilitate and expedite rooftop solar photovoltaic (PV) and battery storage permitting processes. SolarAPP+ allows PV contractors to upload system specifications, have that information automatically reviewed for code compliance, and receive instant approval for code-compliant systems, reducing authority having jurisdiction (AHJ) staff time needed for review. SolarAPP+ also provides inspection checklists to verify installation practices and adherence to approved designs. This report is part of an ongoing series of reviews of SolarAPP+ performance. Consistent with previous performance reviews, we summarize SolarAPP+ adoption trends to date and compare various metrics for PV systems permitted through SolarAPP+ versus systems permitted through traditional AHJ permitting processes. As of the end of 2024, 799 AHJs had expressed interest in the platform, with 264 fully adopting (215) or piloting (49) the platform. In 2024, 861 installers submitted 37,393 permits through the SolarAPP+ platform, including 27,375 permits for PV+storage systems. SolarAPP+ permits accounted for around 43% of all permits issued in all participating AHJs, and more than 60% of all permits in several participating AHJs. We compare permitting timelines through SolarAPP+ to traditional AHJ permitting processes to assess the platform's performance. Consistent with previous SolarAPP+ performance reviews, we find that permitting timelines are significantly shorter for SolarAPP+ projects. Based on median timelines, a typical SolarAPP+ project is permitted and inspected 12 business days sooner than traditional projects. We estimate that automatic SolarAPP+ permitting saved around 18,400 hours of AHJ staff time in 2024. Finally, we estimate that SolarAPP+ eliminated over 100,000 business days in permitting-related delays in 2024.

14 SOLAR ENERGY

VISIONARY: Virtual Intelligence System for Optimizing Novel Analytical Research Yields

VISIONARY is an AI system that accelerates energy materials discovery by automatically generating hypotheses about structure-property relationships. It analyzes patterns in materials data, identifies promising correlations, and proposes testable scientific hypotheses without human intervention. By streamlining this reasoning process, VISIONARY helps researchers efficiently identify candidate materials with desired properties, significantly speeding up the materials development pipeline for energy applications. During the project, we developed a standalone application. The application uses a combination of papers provided by the user and data collected from FutureHouse’s dataset to build an understanding of the background that the user wants to explore for the hypothesis.

36 MATERIALS SCIENCE

Data and scripts associated with a manuscript modeling microbial regulation of priming effects

This data package is associated with the publication “Modeling Microbial Regulatory Feedback in Organic Matter Decomposition Identifies Copiotrophic Traits as Key Drivers of Positive Priming” published as a preprint on BioRXiv by Ahamed et al. (2026); https://doi.org/10.1101/2024.08.11.607483. The package contains MATLAB scripts and saved simulation outputs used to implement a cybernetic model of microbial regulation during complex organic matter (OM) decomposition governing priming effects. It includes models of (i) single microbial functional groups (copiotrophic or oligotrophic degraders) and (ii) binary consortia composed of degraders and non-degraders with contrasting or common growth traits. Simulation results were generated using Monte Carlo analyses, with randomized key model parameters across a range of environmental mixing fractions of complex and labile OM. The dataset was created to provide a transparent and reusable computational framework for systematically exploring how microbial growth traits, metabolic regulation, and community composition influence OM decomposition dynamics and priming effects. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes the variable definitions. This package includes: (1) annotated MATLAB code implementing the system of ordinary differential equations and cybernetic control laws; (2) saved output files containing data (e.g., biomass, substrates, enzyme levels, priming metrics); and (3) scripts for processing saved outputs and regenerating figures. Specifically, the data package contains three main MATLAB scripts: runPrimingModel.m, runPlotData.m, and runPlotSuppFigS1.m, along with this readme and supporting documentation. Users should begin with runPrimingModel.m, which contains the annotated code implementing the system of ordinary differential equations and cybernetic control laws. This script runs the Monte Carlo simulations of microbial OM decomposition and allows users to modify microbial trait definitions, adjust parameter distributions, or define new community configurations. Simulation outputs are automatically saved as .mat files in the folder named SavedData, which stores all pre-generated results included in this package. The second script, runPlotData.m, reads files from the SavedData folder and processes them to regenerate the figures presented in the manuscript. The third script, runPlotSuppFigS1.m, specifically generates Figure S1 in the Supplementary Material of the manuscript. The package also includes the aforementioned files in non-proprietary .txt format. If users intend to use them, they should first save the files in their respective .m or .mat formats prior to execution in MATLAB.

Biomass concentration

SEED: Semantic Energy Exploration and Discovery

The Bioenergy Knowledge Discovery Framework (KDF) hosts a vast repository of specialized data, yet traditional keyword-based search methods often struggle to provide direct answers, requiring significant domain expertise and manual effort to filter through raw documents. To overcome these barriers, this software introduces a semantic search engine that enables both specialists and non-specialists to query the KDF using natural language. By shifting from rigid keyword matching to intent-based retrieval, the tool automatically identifies and ranks the most relevant sources within the database. The system functions by processing natural language queries to extract the most pertinent information, delivering an AI-generated plain-language summary alongside exact supporting quotes from retrieved documents. This integrated approach provides users with immediate, evidence-based answers while eliminating the need for exhaustive manual review. By surfacing direct insights and contextual evidence, the software enhances the usability of existing KDF resources and democratizes access to complex bioenergy data. Ultimately, this semantic search solution accelerates the discovery process and supports faster, more informed decision-making across the bioenergy sector.

Pan, Meiyu (Melrose) [Oak Ridge National Laborator

A semi–automatic analytical methodology for characterizing the energy consumption of MRI systems using load duration curves

Background and purpose: Magnetic resonance imaging (MRI) scanners are a major contributor to greenhouse gas emissions from the healthcare sector, and efforts to improve energy efficiency and reduce energy consumption rely on quantification of the characteristics of energy consumption. The purpose of this work was to develop a semi-automatic analytical methodology for the characterization of the energy consumption of MRI systems using only the load duration curve (LDC). LDCs are a fundamental tool used across various fields to analyze and understand the behavior of loads over time. Methods: An electric current transformer sensor and data logger were installed on two 3T MRI scanners from two vendors, termed M1 (outpatient scanner) and M2 (inpatient/emergency scanner). Data was collected for 1 month (7/11/2023 to 8/11/2023). Active power was calculated, assuming a balanced three-phase system, using the average current measured across all three phases, a 480 V reference voltage for both machines, and vendor-provided power factors. An LDC was constructed for each system by sorting the active power values in descending order and computing the cumulative time (in units of percentage) for each data point. The first derivative of the LDC was then computed (LDC’), smoothed by convolution with a window function (sLDC’), and used to detect transitions between different system modes including (in descending power levels): scan, prepared-to-scan, idle, low-power, and off. The final, segmented LDC was used to measure time (% total time), total energy (kWh), and mean power (kW) for each system mode on both scanners. The method was validated by comparing mean power values, computed using the segmented 1-month LDC, for each nonproductive system mode (i.e., prepared-to-scan, idle, lower-power, and off) against power levels measured after a deliberate system shutdown was performed for each scanner (1 day worth of data). Results: The validation revealed differences in mean power values <1.4% for all nonproductive modes and both scanners. In the scan system mode, the mean power values ranged from 29.8 to 37.2 kW and the total energy consumed for 1 month ranged from 11 106 to 14 466 kWh depending on the scanner. Over the course of 1 month, the portion of time the scanners were in nonproductive modes ranged from 76% to 80% across scanners and the nonproductive energy consumption ranged from 8010 to 6722 kWh depending on the scanner. The M1 (outpatient) scanner consumed 99.9 and 183.9 kWh/day in idle mode for weekdays and weekends, respectively, because the scanner spent 23% more time proportionally in idle mode on the weekends. Conclusions: A semi-automatic method for quantifying energy consumption characteristics of MRI scanners was introduced and validated. This method is relatively simple to implement as it requires only power data from the scanners and avoids the technical challenges associated with extracting and processing scanner log files. Finally, the methodology enables quantitative evaluation of the power, time, and energy characteristics of MRI scanners in scan and nonproductive system modes, providing baseline data and the capability of identifying potential opportunities for enhancing the energy efficiency of MRI scanners.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Generalizable Web User Interface for Scalable and Streamlined Deployment of Building Energy Management Systems in Small and Medium-Sized Commercial Buildings

Small and medium-sized commercial buildings (SMCBs) comprise 94% of US commercial buildings yet face significant barriers to implementing building energy management systems despite advances in smart device technology. Existing solutions present critical limitations: cloud-based API solutions simplify deployment but create vendor lock-in constraints; commercial integrated software solutions ensure compatibility via standardized protocols but require substantial cost and technical expertise; open-source IoT platforms offer cost-effective vendor independence but provide insufficient standardized protocol support for commercial building automation. This research presents a generalizable web user interface framework that bridges the gap between evolving smart device capabilities and lagging software infrastructure for SMCBs. The proposed system integrates VOLTTRON open-source middleware with an automated configuration converter that transforms unified specifications written in YAML, a human-readable data-serialization format, into system-specific files, streamlining manual setup processes. The vendor-agnostic architecture supports industry-standard protocols (BACnet and Modbus) and semantic building models while providing adaptive web interfaces that dynamically adjust to various building configurations. Demonstrations through simulation-based testing and a field deployment show automatic interface adaptation across heterogeneous HVAC systems and multizone monitoring. The automated configuration converter also substantially reduces labor-intensive setup.

Chung, Jihoon [ORNL] (ORCID:0000000184880815)

MAPSTER: Automated Geospatial Data Sharing – Version 1.4.0

The US Department of Energy’s (DOE) Oak Ridge National Laboratory (ORNL) developed MAPSTER which is a geospatial data management tool that aggregates, organizes, and shares data from dispersed sources such as unmanned aerial systems (UAS). Built specifically for use in environments where communications may be limited, MAPSTER utilizes two key technologies to effectively manage data in the field and enable easy data sharing with authorized partners: Observer and Checkpoint. Observer is a lightweight software package on an edge device, such as a laptop, that automatically detects newly processed UAS data and sends to a central server called Checkpoint. Checkpoint is a centralized server at ORNL that receives and manages data from all Observer instances. Even in a very low bandwidth environment, Observer can still send information about the UAS data product almost instantly as it generates its own metadata package on the size of KB (kilobytes). MAPSTER is not only for UAS data but for any geospatial data collected at the austere edge and dispersed sources.

97 MATHEMATICS AND COMPUTING

Bayesian stability and force modeling for uncertain machining processes

Accurately simulating machining operations requires knowledge of the cutting force model and system frequency response. However, this data is collected using specialized instruments in an ex-situ manner. Bayesian statistical methods instead learn the system parameters using cutting test data, but to date, these approaches have only considered milling stability. This paper presents a physics-based Bayesian framework which incorporates both spindle power and milling stability. Initial probabilistic descriptions of the system parameters are propagated through a set of physics functions to form probabilistic predictions about the milling process. The system parameters are then updated using automatically selected cutting tests to reduce parameter uncertainty and identify more productive cutting conditions, where spindle power measurements are used to learn the cutting force model. The framework is demonstrated through both numerical and experimental case studies. Results show that the approach accurately identifies both the system natural frequency and cutting force model.

42 ENGINEERING

Semi-automatic image annotation using 3D LiDAR projections and depth camera data

Efficient image annotation is necessary to utilize deep learning object recognition neural networks in nuclear safeguards, such as for the detection and localization of target objects like nuclear material containers (NMCs). This capability can help automate the inventory accounting of different types of NMCs within nuclear storage facilities. The conventional manual annotation process is labor-intensive and time-consuming, hindering the rapid deployment of deep learning models for NMC identifications. This paper introduces a novel semi-automatic method for annotating 2D images of nuclear material containers (NMCs) by combining 3D light detection and ranging (LiDAR) data with color and depth camera images collected from a handheld scan system. The annotation pipeline involves an operator manually marking new target objects on a LiDAR-generated map, and projecting these 3D locations to images, thereby automatically creating annotations from the projections. The semi-automatic approach significantly reduces manual efforts and the expertise in image annotation that is required to perform the task, allowing deep learning models to be trained on-site within a few hours. The paper compares the performance of models trained on datasets annotated through various methods, including semi-automatic, manual, and commercial annotation services. The evaluation demonstrates that the semi-automatic annotation method achieves comparable or superior results, with a mean average precision (mAP) above 0.9, showcasing its efficiency in training object recognition models. Additionally, the paper explores the application of the proposed method to instance segmentation, achieving promising results in detecting multiple types of NMCs in various formations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P

Test and Validate Distributed Coaxial Cable Sensors for in situ Condition Monitoring of Coal-Fired Boiler Tubes

This project aims to test, validate, and advance the technology readiness level (from TRL5 to TRL7) of a novel low-cost distributed stainless-steel/ceramic coaxial cable sensing (SSC-CCS) technology for in situ monitoring of the boiler tube temperature in existing coal-fired power plants. The novel SSC-CCS sensing technology and associated condition-based monitoring (CBM) software to be demonstrated in this project will lead to an improved understanding of the boiler tube failure mechanisms and a prognostic system to improve the overall performance, reliability, and flexibility of the nation’s coal-fired power plant fleet. A boiler tube monitoring system with distributed coaxial cable temperature sensors and a sensor acquisition system was constructed. The high-temperature coaxial cable sensor with a length of 1.3m was made by using a quartz tube (1mm inner diameter (ID) and 6mm outer diameter (OD)) to concentrically separate a 304 stainless-steel (SS) rod (1mm OD) and SS tube (7.94mm OD and 6.16mm ID). The sensor acquisition system includes a vector network analyzer (VNA), a radio frequency (RF) power amplifier, multiple switches and a USB hub. The distributed stainless-steel quartz coaxial cable sensor (SSQ-CCS) had a linear response to temperature with a resolution uncertainty of σ = 0.77℃. To withstand the harsh conditions of 3,300 steam pressures and 800℃ high temperatures, the sensor was shielded by a protective tube made of the same material as the boiler tube. The protection tube had an OD of 1.5 inches and a thickness of 0.25 inches. In the laboratory tests, the sensor showed good sensitivity and fast response. The drift was bounded between +0.33% and -0.67% during a test at 600℃ for 350 hours, indicating good stability of the sensor. A field test was conducted where four sensors were welded on four superheat tubes (SH-Ts) at a coal-fired power station over 400 days. Conventional thermocouples were welded to the superheater tubes alongside the coaxial cable sensors for the purpose of comparison. Two sensors were capable of distributed sensing, with three multiplexed sensing sections. The other two sensors were single section. During the 400-day test period, the power plant experienced startups and shutdowns. At the steady state operations, the temperature of the boiler tube is about 600℃ (1112°F). The sensors recorded the entire coal-firing processes (start-up, steady state, and shut-down) and the glitch event. A GSM modem and a Watchdog were added to the system to ensure reliable data recording. The GSM modem sent daily messages to plant managers and Clemson team to inform the status of the sensor system. If the system was not normally working, the Watchdog would reboot the system automatically. The new coaxial cable based distributed sensing technology has been proven to be successful in both laboratory and field tests. A comprehensive four-stage multi-physics computational framework has been developed to assist the design, optimization, installation, and operation of SSQ-CCS. With the consideration of various operation conditions, we predict the distributions of flue gas temperatures within coal-fired boilers, the temperature correlation between the boiler tube and SSQ-CCS, and the safety of SSQ-CCS. A conditional-based monitoring system is implemented as well. The computational framework developed in this work can guide the future operation of coal-fired plants and other power plants for the safety prediction of boiler operations.

01 COAL, LIGNITE, AND PEAT

SolarSTARTS: Solar-Assisted State-Aware and ResilienT infrastructure System

This final technical report provides a description and results of the design, development, testing, and validation of the Automated Resilience Management System (ARMS) solution, which integrates different systems to collect, store and process relevant information from the power distribution system to automatically manage the operation of multiple energy assets to enhance the resilience of the power grid against high-impact physical incidents (e.g., hurricanes) and cyber threats. The developed solution comprises a suite of novel AI-based algorithms designed to first detect, locate, and classify anomalies by collecting and analyzing data from multiple monitoring and control devices

14 SOLAR ENERGY

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john