facebook

Upstream Gas Processing Facility Eliminated Decision Latency with UptimeAI Agentic Operations Foundation + Rooty AI

 


The Challenge: Simple Questions Took Days to Answer… And Delayed Decisions Were Costing Millions Per Year

At the company s largest gas field, engineering knowledge was spread across documents and systems PI historian, shift logs, work orders, inspection reports, and in the heads of operators with 20-30 years on site. It was difficult for teams to quickly find relevant information at the asset level. For example, a reliability engineer looking for previous examples of abnormal vibration in reciprocating compressors faced a 4-8 hour manual search across systems that weren t built to talk to each other.
The site team relied on manual search and tribal knowledge.  Building a full evidence trail and response plan for the rising vibration event could take 3 days or longer. The timeline was dictated by how quickly newer engineers or operators could locate the experienced operator with the answers to “Has this happened before? When was it? What were the circumstances? What did we do about it?”
The cost wasn t the search time; it was what happened in the time it took them to investigate, decide, and act. Every hour spent hunting for evidence was an hour a degrading compressor kept running, and a 48-hour delay in catching a failure signature meant the difference between a planned lubrication check and a forced outage costing hundreds of thousands of dollars.

The Solution: Contextual Intelligence That Responds Like Seasoned Engineers

The site deployed UptimeAI s Agentic Operations Foundation, a system built to connect their assets, tags, documents, and engineering knowledge AND make that asset level knowledge retrievable in real time plan language queries. Rooty a conversational interface tuned with domain specific skills and expertise is the front end of the foundation layer, designed to answer complex questions with full evidence trail, reasoning, and proof points.

The knowledge graph is built on an ISO 14224 standard hierarchy rather than a bespoke, site specific map of assets and documents. This was critical since the company needed a repeatable, scalable ontology that they could extend to other facilities without rebuilding the logic from scratch. Every inquiry made through Rooty traversed that same graph structure through a live, continuously updated pipeline into document storage, rather than a static snapshot.

Three key capabilities made UptimeAI s Agentic Operations Foundation the obvious choice for this energy company:

  1. Context aware P&ID understanding let Rooty read P&IDs holistically capturing control loops, fail states, and system interdependencies.
  2. Intelligent document processing classified each document by type before metadata was extracted, letting Rooty map serial numbers across documents and resolve incomplete metadata.
  3. Self-updating performance meant the intelligence improved automatically as it was used, requiring no manual tuning or retraining.

The Impact: Decision Gap Closed Before Consequences Materialized

When asked about the abnormal compressor behavior, Rooty retrieved the relevant sensor trends, cross-referenced maintenance history and inspection reports, and responded with a confidence-scored answer and the evidence behind it. The multi-day lag that once came from relying on a single individual’s memory was replaced by a high-fidelity answer available to any engineer, in minutes.

That exchange, repeated across the site’s recurring investigation, troubleshooting, and onboarding questions, reduced the time to decision by >90% and became the basis for a sitewide business case. The earlier a decision was made, the more expensive a consequence it avoided. The site estimated $1M–$3M in annual value from accelerated decision-making — 20–30% from labor savings, the rest from earlier interventions and avoided trips.

 

 

 

 

Maintenance Optimization Agent Identifies >$650K in Upstream Oil & Gas CM & PM Cost Savings

The Challenge: Generic PM Strategies Start and Stay Suboptimal

Across this operator’s production fields, PM strategies for critical rotating equipment were defined at the compressor train level and rarely revisited. Every gas lift compressor train received the same 60-day lube oil analysis regardless of well conditions, run life history, or actual failure data. When reliability engineers did attempt to revisit the strategy, the exercise required weeks of pulling work orders, failure reports, and production data across dozens of wells and compressor skids spread across the field — often manually reconciled between the CMMS and process historian. Given the pace of upstream production operations and the scarcity of reliability engineering time, PM strategy reviews were the first thing to get deprioritized. The company knew that they were losing money from this type of maintenance strategy, but there was too little time and too much inertia to do anything different.

The Solution: Dynamic Optimization, Unique to Every Asset

By automatically evaluating existing PM strategies against current and historical operations, sensor, and work history data, UptimeAI’s Maintenance Optimization Agent overcame the hurdles of reliability engineer time and organizational inertia. The agent mirrored the asset hierarchy — well, skid, train, component — to match the existing work management system, then prioritized assets using Pareto analysis based on maintenance savings opportunity with the highest potential PM & CM cost savings across the field.

For this customer, Gas Lift Compressor Train 3 proved to be the highest-value target. The agent evaluated historical failure and maintenance history using reliability methods such as Weibull and Crow-AMSAA where statistically appropriate, together with operating context and condition data to determine whether failure patterns were wear-out or infant mortality, then generated a ranked set of specific, implementable recommendations — four distinct types in a single view:

  • Increase frequency where the PM-to-CM ratio was out of sync — adding vibration and lube oil checks on cylinders showing early wear signatures to head off costly unplanned trips.
  • Decrease frequency where zero CM events in the window confirmed safe interval extension — recovering technician hours without added risk on low-criticality components.
  • Add new activities where recurring valve and packing failures had no existing PM to address them — auto-drafting the inspection procedure and checkpoints for direct CMMS import.
  • Remove time-based tasks entirely where live sensor data (e.g. rod load, cylinder vibration, cylinder temperature for reciprocating gas lift compressors) confirmed condition-based monitoring was already available — eliminating redundant calendar-driven teardown inspections.

Approved recommendations were synced live to their CMMS (SAP PM), without the need for manual entry. The agent also flagged PMs tied to API/OSHA process safety requirements, locking them from optimization to protect compliance.

The Impact: From 1 Compressor Train to a Field-wide PM Strategy Optimization

After the Maintenance Optimization Agent identified ~$240K in potential savings on a single gas lift compressor train, the operator extended the program across remaining compressor trains and ESP systems field-wide, identifying cost savings of over $650K.

HAZOP Analysis Agent Reduced SME Time Burden by 60% for North American Chemical Manufacturer

The Challenge: HAZOPs Drained Valuable SME Capacity

As required by OSHA, a mid-sized North American chemical producer conducted Hazard & Operability (HAZOP) studies for each unit in the plant on a 5-year frequency. But with 10 process areas and a continually shrinking pool of SMEs to pull in, HAZOP workshops were a major draw on the site’s most experienced personnel.

Every revalidation cycle required process engineers, operators, EHS leaders, and maintenance experts to spend weeks walking through every deviation-guideword row, often searching for missing information or rebuilding process context rather than evaluating risk. The site was seeing a downward trend in key business metrics like product quality, downtime, and maintenance cost. The people they counted on to improve these metrics were trapped in a conference room highlighting P&IDs.

The Solution: Auto-generated HAZOP Drafts Built on Plant Documentation & Expertise

To reduce the preparation burden on SMEs, the company deployed UptimeAI’s HAZOP Analysis Agent to automate the most time-intensive portions of the PHA process. Starting from existing P&IDs, the agent automatically identified equipment, instrumentation, and process flows to build a digital representation of each unit. The agent generated draft HAZOP nodes, organized relevant documentation, and produced initial deviations, causes, consequences, safeguards, and recommendations aligned to the company’s risk framework.

Rather than starting from a blank worksheet, SMEs entered workshops with a fully populated draft analysis already in place. This shifted their role from manually assembling information to validating risk scenarios, refining recommendations, and making higher-value operational decisions.

The Impact: More Time Spent Managing Risk, Less Time Preparing for Reviews

By deploying the HAZOP Analysis Agent, the site significantly reduced the operational burden associated with recurring PHAs.
Hazop Agent - UptimeAI

Results That Scale

After successfully demonstrating the impact of HAZOP Analysis Agent on SME time for one process unit, the site rapidly expanded to include all site process units + sitewide deployments for all other plants under this business unit. The sites are leveraging the unlocked SME capacity for improvement projects targeting process efficiency, reliability, and quality metrics, while the process safety incident rate is at the lowest it has been in the last decade.

Pump Misalignment Diagnosis with Root Cause Agent Prevents $500K Bearing Failure Event

The Challenge

A major global refining company relied on experienced experts to diagnose equipment anomalies detected by their predictive analytics program. But with the number of experienced employees shrinking, response time for these investigations was going up and failures occurring in that dead time between detection, diagnosis, and decision, were becoming more common. When a centrifugal pump at a refining complex began showing elevated vibration in the 2nd stage bearing, all eyes were on the bearing and its lube oil system. But the root cause was going unnoticed.

Multivariate Detection + Full-System Context = Diagnosing Causes, Not Symptoms

UptimeAI Root Cause Agent detected the abnormal vibration when a multivariate model including roughly 30 tags temperatures, pressures, flows, lube oil conditions started to deviate from live vibration values. This was where softwares contribution to alert investigation used to end for the refinery.

Root Cause Agent leveraged context from integration with their CMMS and SharePoint to take an expert like approach to diagnosing the root cause of the sudden increase in vibration. Looking at all available data sources and built in FMEAs, the agent determined the vibration shift was most likely tied to some recent seal work when the unit was returned to service on a temporary cold alignment. Thermal expansion upon return to operation was accelerating bearing wear at a higher than expected rate.

Root Cause Agent issued recommendations to avoid an unplanned bearing failure based on the trajectory of the degradation. By performing a laser alignment with the hot targets, the refinery avoided having to correct a much more serious bearing issue, saving over $500K in maintenance expense and associated unit downtime.

A Repeatable, Confidence-Ranked Root Cause Hypothesis in Minutes

Root Cause Agent assembled a causal chain automatically when the vibration issue was detected. There was no sending the data off to experts to add to their queue to investigate. Instead, the experts were presented with two completely traceable hypotheses. The top carried 90% confidence: thermal growth misalignment following a seal change during the recent turnaround. The causal chain provided full evidence and links to source documentation:

  • n SAP work order flagged that the coupling had been broken apart and aligned only while offline — thermal growth after restart drove the misalignment.
  • A SharePoint search across tens of thousands of unstructured documents surfaced a prior RCA from a sister pump with identical symptoms, plus an OEM troubleshooting guide on pump-to-driver misalignment.
  • The team further refined the hypotheses when they submitted a lube oil lab analysis via “Rooty” AI copilot, which was added as additional evidence that further supported the misalignment hypothesis. The rising iron content across three reports increased diagnostic the maintenance and operations teams confidence further

Results That Scale

After Root Cause Agent successfully diagnosed this misalignment issue, the refiner began leveraging the agent for all their predictive alerts. They saw alert approval rates grew by nearly 20 percentage points, finishing above 70% — against a benchmark in the single digits for their legacy predictive analytics software. The alert approval rating boost indicated a significant increase in alert quality. After many years of generating alert volumes so high only ~5% of alerts could be investigated,

Root Cause Agent was also dramatically lowering the overall number of alerts the team received. Across 15 compressor trains, the platform averaged roughly one alert per asset per month. This high-accuracy, low-volume, completely pre-diagnosed alerting made site-level self-management viable, and kept the people closest to the equipment engaged.

Maintenance Optimization Agent Identifies ~$300K in Coal Power Plant CM & PM Cost Savings

The Challenge: Generic PM Strategies Start and Stay Suboptimal

At one of Asia Pacific’s largest coal power plants, the PM program strategy was defined at the equipment class level and rarely revisited. All ACW pumps got a 90-day oil check. All bearings were replaced every two years. When the site did revisit the strategy, it required weeks of reliability and maintenance expert time. They would manually pull work history records, old failure reports, and relevant process data. There was so much activation energy required to initiate a PM strategy re-evaluation project that it was always put on the backburner.

The Solution: Dynamic Optimization, Unique to Every Asset

By automatically evaluating existing strategies against current and historical operations and work history data UptimeAI s Maintenance Optimization Agent eliminated this activation energy. The agent mirrors the asset hierarchy to match the existing work management system, then paretos out the assets with highest potential PM & CM cost savings.

For this customer, ACW Pump-1A proved to be ripe for PM optimization. The agent analyzed the full PM and CM history using statistical methods like Crow-AMSAA and Weibull to determine whether failure patterns were wear-out or infant mortality, then generated a ranked set of specific, implementable recommendations four distinct types in a single view:

  •  Increase frequency where the PM-to-CM ratio was out of sync doing more preventative work to reduce corrective costs that outweigh the PM spend increase.
  • Decrease frequency where zero CM events in the window confirmed safe interval extension, recovering maintenance hours without added risk.
  • Add new activities where recurring failures had no existing PM to address them auto-drafting the procedure and checkpoints for direct CMMS import.
  • Remove time-based tasks entirely where live sensor data confirmed condition-based monitoring was already available, eliminating redundant calendar-driven work.

Approved recommendations were synched live to their CMMS (SAP), without the need for manual entry. The agent also flagged PMs that were regulatory requirements locking them from the optimization.

Image: A pareto chart shows the ranked cost savings of the power plant unit operations.

The Impact: From 1 ACW Pump to a Site-wide PM Strategy Optimization

After Maintenance Optimization Agent identified ~$100k in potential savings across a single ACW pump, the site was eager to see where else they could make a dent in their maintenance expenses. The program was further extended identifying cost savings of nearly $300k

 

Leading Cement Enterprise Prevents $50k Coal Mill Bag House Failure by Unlocking Expert Decision Capacity

Leading US Cement Producer Saves $500K in Avoided Issues Caught in the First 4 Weeks After Deploying UptimeAI Reasoning Engine

In large cement operations, kiln availability defines plant performance.
The main drive motor runs continuously under extreme thermal and mechanical loads, where even gradual bearing degradation can trigger an unplanned shutdown and halt clinker production across thousands of tons per day.

For this cement producer, the challenge wasn’t lack of data or alarms, as they already alert monitoring solutions.
It was identifying which subtle deviations actually mattered early enough to act, before they escalated into forced downtime.

UptimeAI’s AI Reasoning Agent continuously reasoned across motor behavior, lubrication performance, and historical failure patterns to detect a developing risk that conventional monitoring systems would have treated as normal variation.

  • By reasoning across asset context, historical patterns, and failure-mode knowledge, the system identified a developing lubrication-related risk well before traditional thresholds were crossed.

Guided by expert-grade recommendations, the plant prepared corrective action during a planned outage, avoiding a forced shutdown. The implementation was quite successful as this leading US Cement Producer saved

$500K in avoided issues caught in the first 4 weeks after deploying UptimeAI Reasoning Engine and a potential ~$550k in production impact.

 

Get a copy of the Customer Story: Download here

 

Preventing NDE Bearing Failure: How a 2 MW Turbine Avoided 169 MWh of Clean Energy Loss with AI Expert

In wind energy operations, the reliability of the non-drive-end (NDE) bearing is fundamental to ensuring rotor stability at ~1500 RPM and safeguarding the generator shaft and drivetrain. Even slight bearing degradation can disrupt generator speed and power stability, threatening both asset availability and revenue generation in highly competitive renewable markets.

In this case, the turbine began showing subtle early deviations, such as the bearing temperature climbing above 75 °C, spikes in vibration acceleration, and correlated dips in generator speed and reactive power. These minor anomalies, often missed by conventional predictive analytics solutions, were immediately detected by the UptimeAI Expert platform:

  • Rising NDE bearing temperature above advisory threshold
  • Vibration increase to nearly 1.8 mm/s
  • Correlation with abnormal generator speed and power swings

Learn how the plant team, guided by UptimeAI’s prescriptive recommendations, were able to connect the dots to identify roller-element wear weeks in advance. Then, during a scheduled downtime event, the team planned the bearing replacement, and ultimately avoided 169 MWh of lost clean energy and 151 hours of downtime

 

Get a copy of the Case Study: Download Here

 

UptimeAI Prevents Gearbox Failure in Vertical Roller Mill Unit and Saves $500K in Production Loss

In cement operations, the reliability of the Vertical Roller Mill (VRM) underpins up to 30% of grinding efficiency and accounts for approximately 25% of plant energy usage. VRMs also enable significant cost savings by reducing energy consumption compared to ball mills.

Hence, ensuring the operational reliability of this asset is critical to achieving the target cost per ton of cement production. Now, even a 24-hour production halt can result in a production loss of $500K (10,000 TPD Ă— $50/ton).

In this case, the cement plant had an output of 10,000 TPD; hence, the role of the VRM unit was essential to achieve this output.
The VRM unit encountered 3 critical anomalies, which were detected on the UptimeAI Expert solution:

  • Steady decline in gearbox oil pressure
  • Rise in thrust-pad temperature
  • Correlated with abnormal lubrication behaviour

Learn how the site team followed UptimeAI’s prescriptive recommendation to prevent a lubrication system failure in this critical asset, restoring optimal flow one week early to avoid production loss.

 

Get a copy of the Case Study: Download Here

 

UptimeAI Improves Asset Reliability & Performance in Thermal Power Plant

Fuel costs account for 60-70% of a power plant’s total operating expenses. Also according to the EIA, the industry benchmark for heat rate is 7,500-8,500 Btu/kWh; however, inefficiencies can increase fuel use by 2-4%, resulting in annual costs of millions of dollars. Here, UptimeAI Expert can work as a unified AI-driven platform to bridge the gaps between reactive maintenance and true operational excellence.

In this case, the thermal power plant producer encountered multiple operational challenges, and the UptimeAI Expert platform was shortlisted to improve asset reliability across 3 phases..

Challenges encountered by the Power Plant:

  • 👉 Anomaly Detection Without Context:
    Anomalies were detected based on DCS alarms without incorporating domain knowledge, relying only on threshold parameters, triggering multiple false alarms and affecting power generation.
  • 👉 Lack of Asset Failure History for Critical Assets:
    No record of equipment failures or corrective actions taken in the past to guide maintenance decisions, which could save time for future performance.
  • 👉 Inefficient Collaboration across the Reliability Team:
    Difficulty tracking actual issues and collaborating across departments due to the use of multiple disconnected applications for critical alerts

Discover how UptimeAI helped a 3x 660 MW thermal power producer overcome operational challenges and achieve $10 million in annual operational savings.

 

Get a copy of the Customer Story: Download Here