facebook

The “Yesterday” Problem: What I’ve learned from customers who challenged the traditional procurement timeline  

By Jag Gattu, CEO, UptimeAI  

This blog was originally published on LinkedIn.

 

I have this question I ask almost every operations leader I meet: when do you actually need this running? I ask it half out of curiosity and half because I already know the answer. Nine times out of ten, it’s some version of “yesterday.” Somebody’s losing sleep over a unit that keeps tripping, or a maintenance backlog that never shrinks, or a retirement that’s about to walk two decades of institutional memory out the door. There’s real urgency in the room.  

And then, almost every time, that urgency hits a wall when it comes to the procurement process. First, they’re waiting on legal. Next, it’s IT. Someone has to find out who actually owns all the systems that need to be connected together. And by the time all of that has happened in sequence, the six months that everyone swore they didn’t have at the start of the project have quietly gone by.  

I want to be careful here, because I don’t think the lesson is “move faster at all costs.” I’ve watched companies get burned by exactly that instinct — skipping a scoping session or waving off a data question can add more delays later. Whatever we build, it has to earn its way into a plant’s operations, and that takes real diligence. What I’ve learned instead, from watching a lot of these cycles up close, is that diligence doesn’t have to walk a straight line. Many procurement tasks can occur simultaneously; they just usually don’t because it’s not the way things have historically been done.   

Here are a few of my favorite techniques that our customers have used to accelerate the value generation of their AI project by choosing not to wait on traditional timelines.   

1. Be an advocate for imperfect data  

I remember a call with an operations director at a cement plant who was clearly ready to move, but his team kept getting stuck arguing over whether they had “enough” and “good enough” data. Every week someone identified a new handful or tags, or a new sensor that would be “nice to have when trying to predict XYZ.” It could have gone on indefinitely.  

At some point he just made a call: they’d agree on the handful of failure modes and the data around the assets that actually mattered for proving the case, then treat everything else as things that could be done in parallel once the solution was live. That one decision point—to accept that some things would be ready now, others would improve over time, and even more opportunities would be uncovered by their use of the technology—avoided weeks of delays in getting the software operational.  

2. Know your scope, your systems, and your stakeholders  

We worked with a project manager for deployment at a large refining client whose team showed up on day one with a clearly defined pilot scope, the KPIs they’d measure success on, the data they’d need to deliver successful results, and the people with the right access to the relevant systems to make it all happen.   

This might sound trivial, but I saw another engagement get delayed by over a month because a new data source was added to the scope which lived at a different level of the enterprise architecture and had a different owner than our other refinery data sources. That owner wasn’t up to speed on our project, and we were fighting for their time with other priority projects that they’d known about for much longer.   

3. Identify, and challenge, dependencies 

One of our early power generation customers taught me something I hadn’t fully appreciated: a lot of the delay in these deals comes from waiting on others out of habit rather than requirement. 

Their security and architecture review ran at the same time as the final commercial conversation, not after it, because somebody on their team asked “why would these two things need to happen in order?” and nobody had a good answer. Procurement and legal moved on a similar clock. Nothing about the actual approvals changed — same reviewers, same rigor. What changed was that they finished around the same time the contract did. That business was ready to execute the week they signed, instead of six or eight weeks later, which is close to how long that kind of review usually takes when everyone assumes it must come last.  

What I want you to take away from these stories 

None of these customers moved faster by cutting anything out. They moved faster because they stopped assuming that due diligence must happen in series. Confirming the use case, mapping the data, starting the security and legal review, building the onboarding plan — almost none of it actually depends on each other.   

If your team is telling you they needed this yesterday, I’d take that seriously — not as a reason to skip anything, but as a reason to ask, honestly, which parts of the next few months are actually sequential out of requirement, and which ones are just habit and can be accelerated to help get them the solution they need, yesterday.  

I’m grateful to the customers who have proved to me that it’s possible to go from demo to go-live in less than a quarter. We’re taking those learnings, applying them to every customer sale, and helping our new customers push to the frontier of adoption of AI reasoning agents for industrial operations.   

 

Our Commitment to Customer Value: Why We Asked Someone Else to Grade Our ROI

By Jag Gattu, CEO at UptimeAI 

 

This article originally appeared on LinkedIn.

 

Every industrial software vendor promises transformation: faster diagnoses, fewer surprises, a healthier bottom line.  

Today’s market places too much responsibility on the customer when it comes to validating these promises. For executives deciding where to put scarce digital investment dollars, the ambiguity gets expensive, and it can be the reason that a promising pilot never reaches enterprise scale. It’s not because technology isn’t valuable, but you need to be able to build a business case that finance can trust. 

This is why, in a market swamped with self-reported ‘proven results’ we invested in helping your teams build a confident business case. Conducting economic studies like the Verdantix Verified Value Delivery (VVD) methodology is not inexpensive and it’s not trivial. It involves hours of interviews with UptimeAI customers to parse out exactly how they are realizing value from the technology, then building that into a calculation framework that can apply across company sizes, industry verticals, and products deployed.  

We think it’s important that our customers understand what value they should expect from the earliest days of engagement – in the form of proof, not promise.  

How do you measure the impact of something that never happened? 

One of the biggest challenges for predictive technologies, particularly in the reliability space, is the definition of success. When our software works exactly as intended, the outcome is often nothing happens. The cost savings of avoiding unplanned failures and unit shutdowns are typically only quantified when the shutdown actually occurs. So how do you quantify the value of something that never happened? 

Since UptimeAI started, our customer success team has worked with every customer to track every alert and diagnosis that prevented a failure and aligned with the customer on what that catch was worth. They comb through past failure events, the lost production, the maintenance expenses, then the system learns the value in warning of that failure mode, so that next time value is automatically assigned. We know that our champions are constantly being asked to justify spend on operational technologies, and we view it as our job to make that task as easy as possible for them.  

Customer interviews confirmed UptimeAI’s commitment to defensible value 

Verdantix conducted interviews of 5 global customers, using their responses to build a model of expected returns for various company demographics. Specifically, for a model $200M-revenue manufacturing site, 3-year ROI of 197% grows to 250% as they expand enterprise-wide, and the software has fully paid for itself in 11 months. In some cases, it’s much faster.  

For one major global cement company – “We had 4 critical alerts in first 6 months of the pilot, when we were very conscious of the value to justify further investment in the solution, and this recouped the cost of the investment.” We continue to compress this payback period by delivering more value, sooner, through more agent products and workflows. 

An exciting future ahead 

197% ROI is exciting, but that’s just getting started. In just 6 months since the study was completed, there are already so many additional value streams that our customers are seeing leveraging new agents. I can’t wait to see what the next year will bring for our current and future customers. 

 

For the full details, read the full Verdantix VVD study.

Build vs. Buy, Part II: The Very Real Barriers to BIY in Highly Regulated Industries

By Jag Gattu, CEO, UptimeAI 

This post originally appeared on LinkedIn.

 

My first post on this topic highlighted the shift we’re seeing in how many tools we build in house vs. buy, and where we draw the line. At UptimeAI, our line sits at the transition point from retrieval to reasoning.  

I took a step back to see where others in the market are drawing the line. While I found general excitement about the potential of BIY (Claude Code and the likes), I found that leaders are being appropriately calculated in their consideration of these projects. There isn’t just one line that they’re drawing for when to build versus buy a solution, but many, and the lines link back to the same challenges that have plagued digital projects in our industry for the past decade.  

Here are the three biggest challenges with BIY agentic AI projects in the process industries.  

1. Governance

In a recent blog post, Verdantix highlighted that AI sprawl is becoming a C-suite problem. Sprawl is a second-order cost to unlimited BIY development. Every vibe-coded agent is also one more tool requiring governance — and a new source of shadow AI, data leakage, and regulatory exposure if it’s not accounted for. Even Amazon has acknowledged internally that “the AI boom is driving duplication and fragmentation, not less of it.” Building a tool and then governing it responsibly, indefinitely, in a live plant are two different jobs, and the governance doesn’t get easier just because the build got faster. Industrial software companies have been operating in this area for decades and have the data governance and cybersecurity infrastructure already in place to deploy agentic AI projects fast without risk. 

2. Scale

According to McKinsey’s 2025 State of AI report, less than 10% of organizations currently experimenting with building AI agents are scaling them. BIY pilot development is easier than production deployment. You can have a demo working reliably on clean, well-behaved offline data in a matter of hours. Standing up something that holds up to the challenges of real operations data — fragmented, inconsistently modeled ERP, MES, historian, and SCADA data, spanning reliability, process, and operations teams — is another. Agents built on top of fragmented data don’t fix fragmented decision-making, they automate and amplify it. That’s true whether you built the agent or bought it. The difference is whether you’re building the curation, integration, and verification layer from scratch, or standing on one that’s already been proven across other plants like yours. 

3. Impact

In a recent LNS Research blog post, Research Analyst Vivek Murugesan emphasizes the importance of “leading with business problems and not technology itself.” Are you looking for a personal productivity tool or a way to make margin-impacting operations decisions? A tool that surfaces the right document or summarizes a work order can be easily built, and it is useful, but it isn’t the same as a system that provides operators with confident decisions to high-stakes calls. Margin is lifted through improvements to availability, quality, and reliability. Personal productivity can support improvements to these areas, but it won’t move the needle on margin the way the components of operational excellence can. An enterprise deployment of a Root Cause Agent, for example, requires automated expert reasoning, robust data handling, and continuous learning and improvement loop. That’s not retrieval, it’s judgment, and it’s proven to provide rapid technology ROI through improvements to diagnosis time, accuracy, and avoided downtime. 

Draw YOUR right lines 

The challenges raised above aren’t reasons to avoid BIY altogether. There are problems where in-house solutions are sufficient, like our sales enablement tool I mentioned in my last post. Retrieval has become easy to build and scale in-house. But when making build v. buy decisions around reasoning and judgement—the agentic AI capabilities required to make impactful decisions in a regulated industrial environment—trusted partners can effectively overcome the governance, scale, and impact hurdles. 

Oil & Gas Upstream Podcast: How AI Reasoning Agents Improve Operational Decision-Making

In this recent episode of the Oil & Gas Upstream podcast (part of Oil & Gas Global Network), host and former US DOE director Elena Melchert interviewed UptimeAI CEO Jag Gattu. They discuss the accelerating expertise drop off in the oil and gas industry, and the need to capture that while managing increasingly complex operations. Jag explains how AI Reasoning Agents are helping engineers make faster, context-aware decisions by combining operational data sources with domain skills and engineering expertise.

Listen to the full episode below.

Beyond the Conversation:

AI Reasoning Agents Make a Big Impact in Upstream Oil & Gas

Read how an upstream oil & gas operator shortened decision cycles by 90% and achieved up to $3 million in annual savings using UptimeAI’s AI Reasoning Agents.

Read the Reinjection Case Study

The Stories Behind the ROI: What 197% Looks Like on the Plant Floor

By Shruti Kela, Engagement Manager at UptimeAI 

 

This blog post originated on LinkedIn.

 

I have spent my career on one question from both ends: what is the real business impact of the work that keeps a plant running. At SLB, I worked in the field, close enough to the equipment to know what a failure costs when it lands. At McKinsey, I built the business cases that decided whether a plant spent millions to prevent that or lived with the risk. Knowing the price of a breakdown is what makes a prevented one so valuable, and also what makes it so hard to prove. The biggest wins in this business are the failures that never happen, and you cannot show a board a breakdown that was quietly avoided.

Unless you catch it in the act. UptimeAI flags a failure early, names the likely cause, and recommends the fix, so engineers act before the loss lands, and every catch is logged as it happens. The breakdown stays invisible, but the savings do not, and those records are exactly what Verdantix set out to audit when UptimeAI brought them in to conduct a Verified Value Delivery study.

That kind of proof is rare. In their 2026 Global Corporate Survey, Verdantix found that 85% of industrial firms say measuring the ROI of their AI tools is a real barrier, which is how promising pilots quietly die. Working from the records of five UptimeAI customers, their independent study landed on a 197% three-year return for a typical site, growing toward 250% at scale.

A high-level metric like that is only as trustworthy as the data that sits under it, so instead of the top-down math, I sat down with my colleagues on UptimeAI’s Customer Success team to sum up the individual moments that compound into triple-digit ROI.

The building blocks for triple-digit ROI

A failure with no alarm

At a cement plant, the kiln is the whole business. Every ton of clinker produced runs through it. One of the sneakiest ways to lose output is coating building at the kiln inlet, and there is no alarm for it, because it never shows up as one bad number. It shows up as small shifts across many at once, pressure edging up, gas composition drifting, each too minor to notice alone. Operators cannot watch every point at once. UptimeAI read them as a system, recognized the coating pattern, and flagged it hours ahead with a clear instruction: ease off the fuel, check the chemistry, prepare to clean. The team cleared it during a short, planned stop instead of a reactive multi-day shutdown, and kept the kiln producing. The value of this catch was amplified a few weeks later when the same issue occurred on a kiln at another plant; six figures and dozens of production hours were kept rather than lost.

One flat gauge hiding many moving ones

A gas turbine at a power operation was holding steady at full load, nothing an operator would look at twice. Underneath, UptimeAI saw what no single gauge would: temperatures rising together across the bearings, the wheel space, and the exhaust, all while output stayed flat. Read one by one, nothing alarmed. Read together, they pointed to one cause. UptimeAI traced it to an inlet guide vane that had drifted open, pulling excess air through the compressor and loading the bearings, and pointed the team straight at the vane hardware. They inspected it, found heavy wear, and planned the fix on their own terms instead of losing the machine to a trip. The turbine kept generating the whole way through. Value the customer confirmed: $280,000.

Diagnosing the cause, not the symptom

When a pump at an oil and gas plant started shaking, everyone looked at the bearing. That is where the vibration showed up, so that is where a normal investigation goes. UptimeAI’s Root Cause Agent looked wider. It read dozens of signals together, pulled in the plant’s own maintenance records and years of documents, and built a causal chain in minutes. Its top answer, at 90% confidence, had nothing to do with the bearing. A seal job weeks earlier had been aligned cold, and once the pump warmed up it pulled out of true. The agent even surfaced a write-up on a sister machine with the same story. The team corrected the alignment instead of tearing into the bearing, kept the unit running, and avoided a failure worth more than $500,000. The fix was never where everyone was looking.

When the smartest answer is “it’s not broken”

Not every alert is a real problem, and the safe reaction, shutting down to go look, is exactly how you lose production you never needed to lose. At an oil and gas operator, a bearing alert fired, and the obvious move was to plan a repair and take the unit down. UptimeAI’s Root Cause Agent challenged it, worked through the data, and found the bearing was fine; the instrument reading it was faulty. It cleared several instrument faults in minutes. The team kept the unit running and skipped an outage and a repair it did not need, worth about $750,000.

Right-sizing maintenance, not just cutting it

At a large power plant, the preventative maintenance plan was set by equipment class and almost never revisited. Every pump on the same oil-change clock, every bearing swapped on the same calendar. Revisiting it meant weeks of expert time pulling work histories, so it rarely happened. UptimeAI’s Maintenance Optimization Agent analyzed the plan continuously and provided ranked recommendations to minimize spend across preventative and corrective maintenance. On a single pump, it found four moves at once: do more where failures were slipping through, less where the data proved it safe, add a missing task, and drop a calendar task that live sensors already covered. UptimeAI pushed approved changes straight into the CMMS, no re-keying. One pump surfaced tens of thousands in savings. Extended across the site, close to $300,000, without a reliability engineer touching a spreadsheet.

A capability that sits inside the customer’s team

The real test of a monitoring tool is whether the customer’s own people run it. At a major North American chemicals producer, the engineers approved and closed most alerts themselves, requiring minimal support from the UptimeAI team. When they wanted a second opinion, they asked UptimeAI’s GenAI copilot, Rooty. A pair of catches on a critical pump went onto the maintenance plan before either failed, so the unit kept running. When the team wanted to test UptimeAI’s predictions against their own historical data, it flagged real past failures months ahead of when they actually happened. With UptimeAI assisting them, the site team caught and fixed problems before they ever became problems.

The math underneath

Every one of these stories represents a single line item in the Verdantix calculation of 197% return. The full study lays out where each dollar comes from, an 11-month payback, and how the return climbs toward 250% as the technology is rolled out across plants or sites. If you have ever watched a promising pilot stall because no one could prove what it was worth, a read of the full study is worth your time.

Get the full Verdantix Verified Value Delivery study

Build vs. Buy: Advice for Operations and Digital Executives in the Age of Generative AI

by Jag Gattu, CEO at UptimeAI

 

This article originally appeared on LinkedIn.

 

If you’re a CIO, CDO, of VP of Operations and you aren’t asking which of your SaaS tools could be retired in the age of generative AI, you’re missing a big part of your job right now. 

Here’s a small example from our own experience. Last week, one of our team members used Claude Code to rebuild a sales relationship-mapping tool we’d been paying for as a SaaS subscription. A few days of work, and we no longer need to renew that contract. Genuinely great use of GenAI — a well-scoped, single-purpose application, recreated in-house faster and cheaper than re-negotiating the contract with the vendor. 

Here’s what we didn’t try to rebuild: our CRM. Nobody on our team seriously proposed it, and not because it wasn’t technically possible. It’s because scope, scale, complexity, and the sheer depth of integrations a CRM touches make it a terrible build candidate — even in a world where a single engineer with the right AI tools can do more than an entire team could five years ago. Add in how fast the economics of the big model providers are shifting, and “build it ourselves” gets shakier by the month, not sturdier. 

That contrast is the whole build-versus-buy question in miniature. GenAI has genuinely moved the line on what’s worth building in-house, but it hasn’t erased the line. 

And the data backs this up. MIT’s widely-cited 2025 study (enter “95% of AI pilots fail to make it to production”) on enterprise AI found that internally built AI deployments succeed at roughly half the rate of externally partnered ones — 33% versus 67%. Gartner has separately warned that at least half of generative AI projects will blow through budget due to poor architectural choices, and that most organizations attempting to build custom models will eventually abandon those efforts due to cost, complexity, and technical debt. That’s not a knock on any one company’s engineering team. It’s what happens structurally when you take on a mission-critical or business-critical system that has to keep working, keep learning, and keep being right, indefinitely. 

So where does something like UptimeAI fall on that line? For the majority of industrial organizations, we see a strong case for not building it yourself — but it’s worth being precise about why, because it’s no one reason between scale, complexity, risk, and upside.  

Most organizations that try to BIY this space end up building a knowledge graph, wiring up retrieval over their operational data, and putting a chatbot on top. That’s a real project, and a legitimate one — but it’s retrieval. It answers, “what does the data say.” 

Bridging from retrieval to reasoning is a big leap. Ours come with the data access, the knowledge graph, the contextual relationships between assets and failure modes, the domain skills of experienced engineers, and the orchestration across sub-agents already built in — tuned to respond to specific high-impact business problems like “what’s actually wrong, and what should you do about it” the way your best expert would. That’s not a chatbot with good retrieval. That’s domain expertise and engineering judgment, encoded, and scaled. It’s a much harder thing to stand up yourself, and it’s exactly the layer where buying beats building. 

Retire the SaaS tools GenAI can genuinely replace. Just don’t confuse that win with being able to BIY it all. Reasoning and retrieval are fundamentally different jobs — and at UptimeAI we take pride in doing the hard jobs in a way that generates big returns for your business. 

 

Schedule a demo to see our off-the-shelf reasoning agents for yourself. 

Maintenance Optimization Agent Identifies >$650K in Upstream Oil & Gas CM & PM Cost Savings

The Challenge: Generic PM Strategies Start and Stay Suboptimal

Across this operator’s production fields, PM strategies for critical rotating equipment were defined at the compressor train level and rarely revisited. Every gas lift compressor train received the same 60-day lube oil analysis regardless of well conditions, run life history, or actual failure data. When reliability engineers did attempt to revisit the strategy, the exercise required weeks of pulling work orders, failure reports, and production data across dozens of wells and compressor skids spread across the field — often manually reconciled between the CMMS and process historian. Given the pace of upstream production operations and the scarcity of reliability engineering time, PM strategy reviews were the first thing to get deprioritized. The company knew that they were losing money from this type of maintenance strategy, but there was too little time and too much inertia to do anything different.

The Solution: Dynamic Optimization, Unique to Every Asset

By automatically evaluating existing PM strategies against current and historical operations, sensor, and work history data, UptimeAI’s Maintenance Optimization Agent overcame the hurdles of reliability engineer time and organizational inertia. The agent mirrored the asset hierarchy — well, skid, train, component — to match the existing work management system, then prioritized assets using Pareto analysis based on maintenance savings opportunity with the highest potential PM & CM cost savings across the field.

For this customer, Gas Lift Compressor Train 3 proved to be the highest-value target. The agent evaluated historical failure and maintenance history using reliability methods such as Weibull and Crow-AMSAA where statistically appropriate, together with operating context and condition data to determine whether failure patterns were wear-out or infant mortality, then generated a ranked set of specific, implementable recommendations — four distinct types in a single view:

  • Increase frequency where the PM-to-CM ratio was out of sync — adding vibration and lube oil checks on cylinders showing early wear signatures to head off costly unplanned trips.
  • Decrease frequency where zero CM events in the window confirmed safe interval extension — recovering technician hours without added risk on low-criticality components.
  • Add new activities where recurring valve and packing failures had no existing PM to address them — auto-drafting the inspection procedure and checkpoints for direct CMMS import.
  • Remove time-based tasks entirely where live sensor data (e.g. rod load, cylinder vibration, cylinder temperature for reciprocating gas lift compressors) confirmed condition-based monitoring was already available — eliminating redundant calendar-driven teardown inspections.

Approved recommendations were synced live to their CMMS (SAP PM), without the need for manual entry. The agent also flagged PMs tied to API/OSHA process safety requirements, locking them from optimization to protect compliance.

The Impact: From 1 Compressor Train to a Field-wide PM Strategy Optimization

After the Maintenance Optimization Agent identified ~$240K in potential savings on a single gas lift compressor train, the operator extended the program across remaining compressor trains and ESP systems field-wide, identifying cost savings of over $650K.

Going Beyond Chatbots: Choosing the Right AI Problems in Oil & Gas

By Jagadish Gattu, CEO at UptimeAI

This article originally appeared on LinkedIn.

The AI race is on, and a pattern I keep seeing across the industry is worth examining. Executives are watching what retailers, banks, and healthcare companies are doing with AI and asking their teams to replicate it. The result is a growing number of chatbots, internal Q&A tools, and productivity assistants that help engineers find documents faster or summarize meeting notes. 

The instinct to move quickly is the right one. What’s worth examining is whether the problems these tools are solving are the highest value ones. 

Matching the Tool to the Problem 

It’s easy to see why consumer AI deployments look appealing. Large language models handle customer service queries, summarize documents, and answer employee questions. These tools are visible, fast to deploy, and easy to demo. For high-volume, repetitive interactions, they deliver genuine value. 

But for leaders in process and asset-intensive industries, the relevant question isn’t whether these tools work. It’s whether they’re addressing problems that actually move the needle for your business. 

Competitive advantage in this industry doesn’t live in a marketing funnel or a customer service queue. It lives in the judgment call that a 30-year veteran makes when a pressure reading doesn’t look right. It lives in the difference between a planned shutdown and an unplanned one that costs $5–30 million a day. 

The Real Challenge in Oil & Gas 

The defining challenge facing process industries today is not productivity. It is expertise at scale. We hear it from customers consistently: decades of institutional knowledge are walking out the door as experienced engineers retire. And the decisions that protect asset integrity, optimize production, and preserve profitability are not simple lookups. They require synthesizing P&IDs, equipment criticality assessments, maintenance history, operating conditions, regulatory constraints, and capital planning frameworks — simultaneously, in context, under pressure, then making judgment calls. 

General-purpose large language models were not designed for this. They weren’t built to interpret a P&ID, assess the criticality of a failing valve within a specific process configuration, or recommend a run-to-failure versus intervention decision based on constantly changing product economics and feed cost structures. They retrieve and summarize well. But the problems that matter most in oil and gas require a different kind of reasoning — contextual, multi-step, domain-specific. 

Where High-Value AI Actually Lives 

At UptimeAI, the question we push customers toward is not “how do we bring AI to oil and gas?” It is “what are the highest-value, most complex, most expertise-dependent problems in our business, and what kind of AI is actually capable of solving them?” 

In upstream and midstream operations, those problems tend to look like: How do you make optimal capital allocation decisions across a portfolio of aging assets with incomplete data? How do you predict the cascading effects of an equipment failure before it happens? How do you encode the judgment of your best engineers and make it available to every operator on every shift? 

These challenges require AI systems capable of multi-step reasoning — systems that can ingest operational data, apply domain-specific logic, weigh tradeoffs, and produce optimal decisions that are explainable and actionable. This is where reasoning agents become genuinely valuable: not as a novelty, but as a multiplier on the expertise your business already has, making it available faster, at greater scale, and with greater consistency than any human team could sustain alone. The key to solving the hardest problems has never been retrieval. From the earliest days of industrial operations, it has always been reasoning.  

A Different Standard for Strategic AI 

Before committing to an AI initiative, it’s worth stress-testing it against three questions. 

1. Does this solve a problem that is genuinely high-value for our business, or are we replicating something that worked in a different industry context?  
2. Is this sustainable at scale, or does it require so much human oversight that efficiency gains become difficult to capture?  
3. And will operational teams actually adopt this and integrate it into their daily workflows?  

The last decade offered lessons about digital tools that never scaled to meet business expectations. Those lessons are worth carrying forward. 

Chatbots and other productivity tools have a real role in any organization’s AI strategy. But the companies that will lead in AI-enabled operations are the ones pairing those early wins with investment in systems sophisticated enough to match the complexity of problems unique to process industries. 

That’s where durable operational advantage gets built. 

Pump Misalignment Diagnosis with Root Cause Agent Prevents $500K Bearing Failure Event

The Challenge

A major global refining company relied on experienced experts to diagnose equipment anomalies detected by their predictive analytics program. But with the number of experienced employees shrinking, response time for these investigations was going up and failures occurring in that dead time between detection, diagnosis, and decision, were becoming more common. When a centrifugal pump at a refining complex began showing elevated vibration in the 2nd stage bearing, all eyes were on the bearing and its lube oil system. But the root cause was going unnoticed.

Multivariate Detection + Full-System Context = Diagnosing Causes, Not Symptoms

UptimeAI Root Cause Agent detected the abnormal vibration when a multivariate model including roughly 30 tags temperatures, pressures, flows, lube oil conditions started to deviate from live vibration values. This was where softwares contribution to alert investigation used to end for the refinery.

Root Cause Agent leveraged context from integration with their CMMS and SharePoint to take an expert like approach to diagnosing the root cause of the sudden increase in vibration. Looking at all available data sources and built in FMEAs, the agent determined the vibration shift was most likely tied to some recent seal work when the unit was returned to service on a temporary cold alignment. Thermal expansion upon return to operation was accelerating bearing wear at a higher than expected rate.

Root Cause Agent issued recommendations to avoid an unplanned bearing failure based on the trajectory of the degradation. By performing a laser alignment with the hot targets, the refinery avoided having to correct a much more serious bearing issue, saving over $500K in maintenance expense and associated unit downtime.

A Repeatable, Confidence-Ranked Root Cause Hypothesis in Minutes

Root Cause Agent assembled a causal chain automatically when the vibration issue was detected. There was no sending the data off to experts to add to their queue to investigate. Instead, the experts were presented with two completely traceable hypotheses. The top carried 90% confidence: thermal growth misalignment following a seal change during the recent turnaround. The causal chain provided full evidence and links to source documentation:

  • n SAP work order flagged that the coupling had been broken apart and aligned only while offline — thermal growth after restart drove the misalignment.
  • A SharePoint search across tens of thousands of unstructured documents surfaced a prior RCA from a sister pump with identical symptoms, plus an OEM troubleshooting guide on pump-to-driver misalignment.
  • The team further refined the hypotheses when they submitted a lube oil lab analysis via “Rooty” AI copilot, which was added as additional evidence that further supported the misalignment hypothesis. The rising iron content across three reports increased diagnostic the maintenance and operations teams confidence further

Results That Scale

After Root Cause Agent successfully diagnosed this misalignment issue, the refiner began leveraging the agent for all their predictive alerts. They saw alert approval rates grew by nearly 20 percentage points, finishing above 70% — against a benchmark in the single digits for their legacy predictive analytics software. The alert approval rating boost indicated a significant increase in alert quality. After many years of generating alert volumes so high only ~5% of alerts could be investigated,

Root Cause Agent was also dramatically lowering the overall number of alerts the team received. Across 15 compressor trains, the platform averaged roughly one alert per asset per month. This high-accuracy, low-volume, completely pre-diagnosed alerting made site-level self-management viable, and kept the people closest to the equipment engaged.

UptimeAI is where the market is heading

By Jagadish Gattu, CEO of UptimeAI 

This article first appeared on LinkedIn.

Verdantix biennial Green Quadrant for Asset Performance Management (APM) was recently released, and this version was different. AI was no longer treated as a singular capability but woven into the thread of every evaluation criteria. As the newest company in a market dominated by incumbents we saw this as an advantage. Being born in the age of AI means our AI products were built that way from the ground up, not sprinkled on top of decades old technology. This advantage was reflected in various capability scores, and also the overarching narrative of the report.  

The No. 1 Score in Market Vision & Business Strategy 

The momentum (x) axis in a Green Quadrant correlates strongly to the size of the business and dominance in the marketplace, which can be a big advantage for legacy companies, who have had decades to grow sales and following to where they are today. For newer companies, the place to shine amidst the momentum criteria is less about where you’ve been and more about where you’re going.  

When I saw that UptimeAI had been awarded the highest score in the field for Market Vision & Business Strategy (a 2.9 on a 3.0 scale), I was not surprised. In my article on closing the gap between insight and execution, I discussed the #1 takeaway from the 2025 Verdantix Asset Management Council. The 13 asset management leaders from major energy and industrial companies came together and deduced that: 

Industrial agility – the ability to rapidly adapt operations, processes and workforce focus – is becoming increasingly relevant because of macroeconomic pressure and increasing dislocation. Improving agility by closing the gap between insight and execution will be the key to enhancing operational excellence amidst reskilling, data-fragmentation and scaling challenges. – Verdantix 2025 Asset Management Council

This was a recurring theme in the 2026 Green Quadrant, and a common strength amongst the companies that scored highest in vision and strategy.  

Agentic AI is transforming APM from insights-driven to action-oriented 

Past definitions of APM held up predictive analytics capabilities as the gold standard for uncovering insights hidden in untapped data. But today, the data’s been tapped, the insights are piling up, yet the outcomes still lag. It turns out it was never a shortage of insights limiting our industry’s margins. It is the shortage of expert decisions that drive actions and the subsequent outcomes that we’ve been limited by.

In talks at CERAWeek and other events, I’ve described the expert decision  bottleneck that’s costing industrial organizations millions every year. Predictive analytics stops at the point of detection, awaiting expert interpretation to arrive at a decision. The amount of time spent between uncovering an insight and getting to an optimal decision is the decision latency created by the expert bottleneck. Decision latency costs organizations millions in failures, repairs, unplanned downtime, and excessive preventative maintenance. Overcoming that expert bottleneck is the key to creating an APM program that delivers real margin impact. 

Decision latency in industrial operations
The expert bottleneck created by decades of tools that have focused on insights.

UptimeAI takes an agent-first approach, applying AI to emulate key maintenance workflows such as optimization and root-cause analysis. It continuously evaluates operating conditions, asset criticality, failure mode and effects analysis (FMEA), work orders and equipment documentation to identify underlying failure drivers, optimize maintenance strategies and recommend actions, with transparent reasoning behind each decision. This shift towards agentic AI also raises the bar for incumbents: success is increasingly dependent on either innovating quickly or forming partnerships to effectively leverage agents. – 2026 Verdantix Green Quadrant: APM

When the decision gap closes, results compound 

One of the strengths of the Verdantix research process is their use of customer interviews to validate market presence and product capabilities. Speaking with UptimeAI customers, Verdantix verified maintenance time savings of 70 hours per month and annual reliability + performance savings of $10M per year at a single coal plant for one of India’s top 5 largest power generation companies. In a second interview with a top 10 US cement manufacturer, Verdantix confirmed $500k savings and 24h of avoided kiln downtime from the early diagnosis and mitigation of a single anomaly event.  

When the gap between insights and actionable decisions collapses, margin growth accelerates. UptimeAI has delivered repeatable results across global industry leading organizations in oil and gas, cement, chemicals, and power generation.  

The most forward-looking industrial organizations are already making the shift — and UptimeAI is making it possible.