APAC Webinar: How Agentic AI Delivers Real Results on the Plant Floor
The “Yesterday” Problem: What I’ve learned from customers who challenged the traditional procurement timeline
By Jag Gattu, CEO, UptimeAI
This blog was originally published on LinkedIn.
I have this question I ask almost every operations leader I meet: when do you actually need this running? I ask it half out of curiosity and half because I already know the answer. Nine times out of ten, it’s some version of “yesterday.” Somebody’s losing sleep over a unit that keeps tripping, or a maintenance backlog that never shrinks, or a retirement that’s about to walk two decades of institutional memory out the door. There’s real urgency in the room.
And then, almost every time, that urgency hits a wall when it comes to the procurement process. First, they’re waiting on legal. Next, it’s IT. Someone has to find out who actually owns all the systems that need to be connected together. And by the time all of that has happened in sequence, the six months that everyone swore they didn’t have at the start of the project have quietly gone by.
I want to be careful here, because I don’t think the lesson is “move faster at all costs.” I’ve watched companies get burned by exactly that instinct — skipping a scoping session or waving off a data question can add more delays later. Whatever we build, it has to earn its way into a plant’s operations, and that takes real diligence. What I’ve learned instead, from watching a lot of these cycles up close, is that diligence doesn’t have to walk a straight line. Many procurement tasks can occur simultaneously; they just usually don’t because it’s not the way things have historically been done.
Here are a few of my favorite techniques that our customers have used to accelerate the value generation of their AI project by choosing not to wait on traditional timelines.
1. Be an advocate for imperfect data
I remember a call with an operations director at a cement plant who was clearly ready to move, but his team kept getting stuck arguing over whether they had “enough” and “good enough” data. Every week someone identified a new handful or tags, or a new sensor that would be “nice to have when trying to predict XYZ.” It could have gone on indefinitely.
At some point he just made a call: they’d agree on the handful of failure modes and the data around the assets that actually mattered for proving the case, then treat everything else as things that could be done in parallel once the solution was live. That one decision point—to accept that some things would be ready now, others would improve over time, and even more opportunities would be uncovered by their use of the technology—avoided weeks of delays in getting the software operational.
2. Know your scope, your systems, and your stakeholders
We worked with a project manager for deployment at a large refining client whose team showed up on day one with a clearly defined pilot scope, the KPIs they’d measure success on, the data they’d need to deliver successful results, and the people with the right access to the relevant systems to make it all happen.
This might sound trivial, but I saw another engagement get delayed by over a month because a new data source was added to the scope which lived at a different level of the enterprise architecture and had a different owner than our other refinery data sources. That owner wasn’t up to speed on our project, and we were fighting for their time with other priority projects that they’d known about for much longer.
3. Identify, and challenge, dependencies
One of our early power generation customers taught me something I hadn’t fully appreciated: a lot of the delay in these deals comes from waiting on others out of habit rather than requirement.
Their security and architecture review ran at the same time as the final commercial conversation, not after it, because somebody on their team asked “why would these two things need to happen in order?” and nobody had a good answer. Procurement and legal moved on a similar clock. Nothing about the actual approvals changed — same reviewers, same rigor. What changed was that they finished around the same time the contract did. That business was ready to execute the week they signed, instead of six or eight weeks later, which is close to how long that kind of review usually takes when everyone assumes it must come last.
What I want you to take away from these stories
None of these customers moved faster by cutting anything out. They moved faster because they stopped assuming that due diligence must happen in series. Confirming the use case, mapping the data, starting the security and legal review, building the onboarding plan — almost none of it actually depends on each other.
If your team is telling you they needed this yesterday, I’d take that seriously — not as a reason to skip anything, but as a reason to ask, honestly, which parts of the next few months are actually sequential out of requirement, and which ones are just habit and can be accelerated to help get them the solution they need, yesterday.
I’m grateful to the customers who have proved to me that it’s possible to go from demo to go-live in less than a quarter. We’re taking those learnings, applying them to every customer sale, and helping our new customers push to the frontier of adoption of AI reasoning agents for industrial operations.
Our Commitment to Customer Value: Why We Asked Someone Else to Grade Our ROI
By Jag Gattu, CEO at UptimeAI
This article originally appeared on LinkedIn.
Every industrial software vendor promises transformation: faster diagnoses, fewer surprises, a healthier bottom line.
Today’s market places too much responsibility on the customer when it comes to validating these promises. For executives deciding where to put scarce digital investment dollars, the ambiguity gets expensive, and it can be the reason that a promising pilot never reaches enterprise scale. It’s not because technology isn’t valuable, but you need to be able to build a business case that finance can trust.
This is why, in a market swamped with self-reported ‘proven results’ we invested in helping your teams build a confident business case. Conducting economic studies like the Verdantix Verified Value Delivery (VVD) methodology is not inexpensive and it’s not trivial. It involves hours of interviews with UptimeAI customers to parse out exactly how they are realizing value from the technology, then building that into a calculation framework that can apply across company sizes, industry verticals, and products deployed.
We think it’s important that our customers understand what value they should expect from the earliest days of engagement – in the form of proof, not promise.
How do you measure the impact of something that never happened?
One of the biggest challenges for predictive technologies, particularly in the reliability space, is the definition of success. When our software works exactly as intended, the outcome is often nothing happens. The cost savings of avoiding unplanned failures and unit shutdowns are typically only quantified when the shutdown actually occurs. So how do you quantify the value of something that never happened?
Since UptimeAI started, our customer success team has worked with every customer to track every alert and diagnosis that prevented a failure and aligned with the customer on what that catch was worth. They comb through past failure events, the lost production, the maintenance expenses, then the system learns the value in warning of that failure mode, so that next time value is automatically assigned. We know that our champions are constantly being asked to justify spend on operational technologies, and we view it as our job to make that task as easy as possible for them.
Customer interviews confirmed UptimeAI’s commitment to defensible value
Verdantix conducted interviews of 5 global customers, using their responses to build a model of expected returns for various company demographics. Specifically, for a model $200M-revenue manufacturing site, 3-year ROI of 197% grows to 250% as they expand enterprise-wide, and the software has fully paid for itself in 11 months. In some cases, it’s much faster.
For one major global cement company – “We had 4 critical alerts in first 6 months of the pilot, when we were very conscious of the value to justify further investment in the solution, and this recouped the cost of the investment.” We continue to compress this payback period by delivering more value, sooner, through more agent products and workflows.
An exciting future ahead
197% ROI is exciting, but that’s just getting started. In just 6 months since the study was completed, there are already so many additional value streams that our customers are seeing leveraging new agents. I can’t wait to see what the next year will bring for our current and future customers.
For the full details, read the full Verdantix VVD study.
Build vs. Buy, Part II: The Very Real Barriers to BIY in Highly Regulated Industries
By Jag Gattu, CEO, UptimeAI
This post originally appeared on LinkedIn.
My first post on this topic highlighted the shift we’re seeing in how many tools we build in house vs. buy, and where we draw the line. At UptimeAI, our line sits at the transition point from retrieval to reasoning.
I took a step back to see where others in the market are drawing the line. While I found general excitement about the potential of BIY (Claude Code and the likes), I found that leaders are being appropriately calculated in their consideration of these projects. There isn’t just one line that they’re drawing for when to build versus buy a solution, but many, and the lines link back to the same challenges that have plagued digital projects in our industry for the past decade.
Here are the three biggest challenges with BIY agentic AI projects in the process industries.
1. Governance
In a recent blog post, Verdantix highlighted that AI sprawl is becoming a C-suite problem. Sprawl is a second-order cost to unlimited BIY development. Every vibe-coded agent is also one more tool requiring governance — and a new source of shadow AI, data leakage, and regulatory exposure if it’s not accounted for. Even Amazon has acknowledged internally that “the AI boom is driving duplication and fragmentation, not less of it.” Building a tool and then governing it responsibly, indefinitely, in a live plant are two different jobs, and the governance doesn’t get easier just because the build got faster. Industrial software companies have been operating in this area for decades and have the data governance and cybersecurity infrastructure already in place to deploy agentic AI projects fast without risk.
2. Scale
According to McKinsey’s 2025 State of AI report, less than 10% of organizations currently experimenting with building AI agents are scaling them. BIY pilot development is easier than production deployment. You can have a demo working reliably on clean, well-behaved offline data in a matter of hours. Standing up something that holds up to the challenges of real operations data — fragmented, inconsistently modeled ERP, MES, historian, and SCADA data, spanning reliability, process, and operations teams — is another. Agents built on top of fragmented data don’t fix fragmented decision-making, they automate and amplify it. That’s true whether you built the agent or bought it. The difference is whether you’re building the curation, integration, and verification layer from scratch, or standing on one that’s already been proven across other plants like yours.
3. Impact
In a recent LNS Research blog post, Research Analyst Vivek Murugesan emphasizes the importance of “leading with business problems and not technology itself.” Are you looking for a personal productivity tool or a way to make margin-impacting operations decisions? A tool that surfaces the right document or summarizes a work order can be easily built, and it is useful, but it isn’t the same as a system that provides operators with confident decisions to high-stakes calls. Margin is lifted through improvements to availability, quality, and reliability. Personal productivity can support improvements to these areas, but it won’t move the needle on margin the way the components of operational excellence can. An enterprise deployment of a Root Cause Agent, for example, requires automated expert reasoning, robust data handling, and continuous learning and improvement loop. That’s not retrieval, it’s judgment, and it’s proven to provide rapid technology ROI through improvements to diagnosis time, accuracy, and avoided downtime.
Draw YOUR right lines
The challenges raised above aren’t reasons to avoid BIY altogether. There are problems where in-house solutions are sufficient, like our sales enablement tool I mentioned in my last post. Retrieval has become easy to build and scale in-house. But when making build v. buy decisions around reasoning and judgement—the agentic AI capabilities required to make impactful decisions in a regulated industrial environment—trusted partners can effectively overcome the governance, scale, and impact hurdles.
Oil & Gas Upstream Podcast: How AI Reasoning Agents Improve Operational Decision-Making
In this recent episode of the Oil & Gas Upstream podcast (part of Oil & Gas Global Network), host and former US DOE director Elena Melchert interviewed UptimeAI CEO Jag Gattu. They discuss the accelerating expertise drop off in the oil and gas industry, and the need to capture that while managing increasingly complex operations. Jag explains how AI Reasoning Agents are helping engineers make faster, context-aware decisions by combining operational data sources with domain skills and engineering expertise.
Listen to the full episode below.
Beyond the Conversation:
AI Reasoning Agents Make a Big Impact in Upstream Oil & Gas
Read how an upstream oil & gas operator shortened decision cycles by 90% and achieved up to $3 million in annual savings using UptimeAI’s AI Reasoning Agents.
The Stories Behind the ROI: What 197% Looks Like on the Plant Floor
By Shruti Kela, Engagement Manager at UptimeAI
This blog post originated on LinkedIn.
I have spent my career on one question from both ends: what is the real business impact of the work that keeps a plant running. At SLB, I worked in the field, close enough to the equipment to know what a failure costs when it lands. At McKinsey, I built the business cases that decided whether a plant spent millions to prevent that or lived with the risk. Knowing the price of a breakdown is what makes a prevented one so valuable, and also what makes it so hard to prove. The biggest wins in this business are the failures that never happen, and you cannot show a board a breakdown that was quietly avoided.
Unless you catch it in the act. UptimeAI flags a failure early, names the likely cause, and recommends the fix, so engineers act before the loss lands, and every catch is logged as it happens. The breakdown stays invisible, but the savings do not, and those records are exactly what Verdantix set out to audit when UptimeAI brought them in to conduct a Verified Value Delivery study.
That kind of proof is rare. In their 2026 Global Corporate Survey, Verdantix found that 85% of industrial firms say measuring the ROI of their AI tools is a real barrier, which is how promising pilots quietly die. Working from the records of five UptimeAI customers, their independent study landed on a 197% three-year return for a typical site, growing toward 250% at scale.
A high-level metric like that is only as trustworthy as the data that sits under it, so instead of the top-down math, I sat down with my colleagues on UptimeAI’s Customer Success team to sum up the individual moments that compound into triple-digit ROI.
The building blocks for triple-digit ROI
A failure with no alarm
At a cement plant, the kiln is the whole business. Every ton of clinker produced runs through it. One of the sneakiest ways to lose output is coating building at the kiln inlet, and there is no alarm for it, because it never shows up as one bad number. It shows up as small shifts across many at once, pressure edging up, gas composition drifting, each too minor to notice alone. Operators cannot watch every point at once. UptimeAI read them as a system, recognized the coating pattern, and flagged it hours ahead with a clear instruction: ease off the fuel, check the chemistry, prepare to clean. The team cleared it during a short, planned stop instead of a reactive multi-day shutdown, and kept the kiln producing. The value of this catch was amplified a few weeks later when the same issue occurred on a kiln at another plant; six figures and dozens of production hours were kept rather than lost.
One flat gauge hiding many moving ones
A gas turbine at a power operation was holding steady at full load, nothing an operator would look at twice. Underneath, UptimeAI saw what no single gauge would: temperatures rising together across the bearings, the wheel space, and the exhaust, all while output stayed flat. Read one by one, nothing alarmed. Read together, they pointed to one cause. UptimeAI traced it to an inlet guide vane that had drifted open, pulling excess air through the compressor and loading the bearings, and pointed the team straight at the vane hardware. They inspected it, found heavy wear, and planned the fix on their own terms instead of losing the machine to a trip. The turbine kept generating the whole way through. Value the customer confirmed: $280,000.
Diagnosing the cause, not the symptom
When a pump at an oil and gas plant started shaking, everyone looked at the bearing. That is where the vibration showed up, so that is where a normal investigation goes. UptimeAI’s Root Cause Agent looked wider. It read dozens of signals together, pulled in the plant’s own maintenance records and years of documents, and built a causal chain in minutes. Its top answer, at 90% confidence, had nothing to do with the bearing. A seal job weeks earlier had been aligned cold, and once the pump warmed up it pulled out of true. The agent even surfaced a write-up on a sister machine with the same story. The team corrected the alignment instead of tearing into the bearing, kept the unit running, and avoided a failure worth more than $500,000. The fix was never where everyone was looking.
When the smartest answer is “it’s not broken”
Not every alert is a real problem, and the safe reaction, shutting down to go look, is exactly how you lose production you never needed to lose. At an oil and gas operator, a bearing alert fired, and the obvious move was to plan a repair and take the unit down. UptimeAI’s Root Cause Agent challenged it, worked through the data, and found the bearing was fine; the instrument reading it was faulty. It cleared several instrument faults in minutes. The team kept the unit running and skipped an outage and a repair it did not need, worth about $750,000.
Right-sizing maintenance, not just cutting it
At a large power plant, the preventative maintenance plan was set by equipment class and almost never revisited. Every pump on the same oil-change clock, every bearing swapped on the same calendar. Revisiting it meant weeks of expert time pulling work histories, so it rarely happened. UptimeAI’s Maintenance Optimization Agent analyzed the plan continuously and provided ranked recommendations to minimize spend across preventative and corrective maintenance. On a single pump, it found four moves at once: do more where failures were slipping through, less where the data proved it safe, add a missing task, and drop a calendar task that live sensors already covered. UptimeAI pushed approved changes straight into the CMMS, no re-keying. One pump surfaced tens of thousands in savings. Extended across the site, close to $300,000, without a reliability engineer touching a spreadsheet.
A capability that sits inside the customer’s team
The real test of a monitoring tool is whether the customer’s own people run it. At a major North American chemicals producer, the engineers approved and closed most alerts themselves, requiring minimal support from the UptimeAI team. When they wanted a second opinion, they asked UptimeAI’s GenAI copilot, Rooty. A pair of catches on a critical pump went onto the maintenance plan before either failed, so the unit kept running. When the team wanted to test UptimeAI’s predictions against their own historical data, it flagged real past failures months ahead of when they actually happened. With UptimeAI assisting them, the site team caught and fixed problems before they ever became problems.
The math underneath
Every one of these stories represents a single line item in the Verdantix calculation of 197% return. The full study lays out where each dollar comes from, an 11-month payback, and how the return climbs toward 250% as the technology is rolled out across plants or sites. If you have ever watched a promising pilot stall because no one could prove what it was worth, a read of the full study is worth your time.
Build vs. Buy: Advice for Operations and Digital Executives in the Age of Generative AI
by Jag Gattu, CEO at UptimeAI
This article originally appeared on LinkedIn.
If you’re a CIO, CDO, of VP of Operations and you aren’t asking which of your SaaS tools could be retired in the age of generative AI, you’re missing a big part of your job right now.
Here’s a small example from our own experience. Last week, one of our team members used Claude Code to rebuild a sales relationship-mapping tool we’d been paying for as a SaaS subscription. A few days of work, and we no longer need to renew that contract. Genuinely great use of GenAI — a well-scoped, single-purpose application, recreated in-house faster and cheaper than re-negotiating the contract with the vendor.
Here’s what we didn’t try to rebuild: our CRM. Nobody on our team seriously proposed it, and not because it wasn’t technically possible. It’s because scope, scale, complexity, and the sheer depth of integrations a CRM touches make it a terrible build candidate — even in a world where a single engineer with the right AI tools can do more than an entire team could five years ago. Add in how fast the economics of the big model providers are shifting, and “build it ourselves” gets shakier by the month, not sturdier.
That contrast is the whole build-versus-buy question in miniature. GenAI has genuinely moved the line on what’s worth building in-house, but it hasn’t erased the line.
And the data backs this up. MIT’s widely-cited 2025 study (enter “95% of AI pilots fail to make it to production”) on enterprise AI found that internally built AI deployments succeed at roughly half the rate of externally partnered ones — 33% versus 67%. Gartner has separately warned that at least half of generative AI projects will blow through budget due to poor architectural choices, and that most organizations attempting to build custom models will eventually abandon those efforts due to cost, complexity, and technical debt. That’s not a knock on any one company’s engineering team. It’s what happens structurally when you take on a mission-critical or business-critical system that has to keep working, keep learning, and keep being right, indefinitely.
So where does something like UptimeAI fall on that line? For the majority of industrial organizations, we see a strong case for not building it yourself — but it’s worth being precise about why, because it’s no one reason between scale, complexity, risk, and upside.
Most organizations that try to BIY this space end up building a knowledge graph, wiring up retrieval over their operational data, and putting a chatbot on top. That’s a real project, and a legitimate one — but it’s retrieval. It answers, “what does the data say.”
Bridging from retrieval to reasoning is a big leap. Ours come with the data access, the knowledge graph, the contextual relationships between assets and failure modes, the domain skills of experienced engineers, and the orchestration across sub-agents already built in — tuned to respond to specific high-impact business problems like “what’s actually wrong, and what should you do about it” the way your best expert would. That’s not a chatbot with good retrieval. That’s domain expertise and engineering judgment, encoded, and scaled. It’s a much harder thing to stand up yourself, and it’s exactly the layer where buying beats building.
Retire the SaaS tools GenAI can genuinely replace. Just don’t confuse that win with being able to BIY it all. Reasoning and retrieval are fundamentally different jobs — and at UptimeAI we take pride in doing the hard jobs in a way that generates big returns for your business.
Schedule a demo to see our off-the-shelf reasoning agents for yourself.
Upstream Gas Processing Facility Eliminated Decision Latency with UptimeAI Agentic Operations Foundation + Rooty AI

The Challenge: Simple Questions Took Days to Answer… And Delayed Decisions Were Costing Millions Per Year
At the company s largest gas field, engineering knowledge was spread across documents and systems PI historian, shift logs, work orders, inspection reports, and in the heads of operators with 20-30 years on site. It was difficult for teams to quickly find relevant information at the asset level. For example, a reliability engineer looking for previous examples of abnormal vibration in reciprocating compressors faced a 4-8 hour manual search across systems that weren t built to talk to each other.
The site team relied on manual search and tribal knowledge. Building a full evidence trail and response plan for the rising vibration event could take 3 days or longer. The timeline was dictated by how quickly newer engineers or operators could locate the experienced operator with the answers to “Has this happened before? When was it? What were the circumstances? What did we do about it?”
The cost wasn t the search time; it was what happened in the time it took them to investigate, decide, and act. Every hour spent hunting for evidence was an hour a degrading compressor kept running, and a 48-hour delay in catching a failure signature meant the difference between a planned lubrication check and a forced outage costing hundreds of thousands of dollars.
The Solution: Contextual Intelligence That Responds Like Seasoned Engineers
The site deployed UptimeAI s Agentic Operations Foundation, a system built to connect their assets, tags, documents, and engineering knowledge AND make that asset level knowledge retrievable in real time plan language queries. Rooty a conversational interface tuned with domain specific skills and expertise is the front end of the foundation layer, designed to answer complex questions with full evidence trail, reasoning, and proof points.
The knowledge graph is built on an ISO 14224 standard hierarchy rather than a bespoke, site specific map of assets and documents. This was critical since the company needed a repeatable, scalable ontology that they could extend to other facilities without rebuilding the logic from scratch. Every inquiry made through Rooty traversed that same graph structure through a live, continuously updated pipeline into document storage, rather than a static snapshot.
Three key capabilities made UptimeAI s Agentic Operations Foundation the obvious choice for this energy company:
- Context aware P&ID understanding let Rooty read P&IDs holistically capturing control loops, fail states, and system interdependencies.
- Intelligent document processing classified each document by type before metadata was extracted, letting Rooty map serial numbers across documents and resolve incomplete metadata.
- Self-updating performance meant the intelligence improved automatically as it was used, requiring no manual tuning or retraining.

The Impact: Decision Gap Closed Before Consequences Materialized
When asked about the abnormal compressor behavior, Rooty retrieved the relevant sensor trends, cross-referenced maintenance history and inspection reports, and responded with a confidence-scored answer and the evidence behind it. The multi-day lag that once came from relying on a single individual’s memory was replaced by a high-fidelity answer available to any engineer, in minutes.
That exchange, repeated across the site’s recurring investigation, troubleshooting, and onboarding questions, reduced the time to decision by >90% and became the basis for a sitewide business case. The earlier a decision was made, the more expensive a consequence it avoided. The site estimated $1M–$3M in annual value from accelerated decision-making — 20–30% from labor savings, the rest from earlier interventions and avoided trips.

Maintenance Optimization Agent Identifies >$650K in Upstream Oil & Gas CM & PM Cost Savings

The Challenge: Generic PM Strategies Start and Stay Suboptimal
Across this operator’s production fields, PM strategies for critical rotating equipment were defined at the compressor train level and rarely revisited. Every gas lift compressor train received the same 60-day lube oil analysis regardless of well conditions, run life history, or actual failure data. When reliability engineers did attempt to revisit the strategy, the exercise required weeks of pulling work orders, failure reports, and production data across dozens of wells and compressor skids spread across the field — often manually reconciled between the CMMS and process historian. Given the pace of upstream production operations and the scarcity of reliability engineering time, PM strategy reviews were the first thing to get deprioritized. The company knew that they were losing money from this type of maintenance strategy, but there was too little time and too much inertia to do anything different.
The Solution: Dynamic Optimization, Unique to Every Asset
By automatically evaluating existing PM strategies against current and historical operations, sensor, and work history data, UptimeAI’s Maintenance Optimization Agent overcame the hurdles of reliability engineer time and organizational inertia. The agent mirrored the asset hierarchy — well, skid, train, component — to match the existing work management system, then prioritized assets using Pareto analysis based on maintenance savings opportunity with the highest potential PM & CM cost savings across the field.
For this customer, Gas Lift Compressor Train 3 proved to be the highest-value target. The agent evaluated historical failure and maintenance history using reliability methods such as Weibull and Crow-AMSAA where statistically appropriate, together with operating context and condition data to determine whether failure patterns were wear-out or infant mortality, then generated a ranked set of specific, implementable recommendations — four distinct types in a single view:
- Increase frequency where the PM-to-CM ratio was out of sync — adding vibration and lube oil checks on cylinders showing early wear signatures to head off costly unplanned trips.
- Decrease frequency where zero CM events in the window confirmed safe interval extension — recovering technician hours without added risk on low-criticality components.
- Add new activities where recurring valve and packing failures had no existing PM to address them — auto-drafting the inspection procedure and checkpoints for direct CMMS import.
- Remove time-based tasks entirely where live sensor data (e.g. rod load, cylinder vibration, cylinder temperature for reciprocating gas lift compressors) confirmed condition-based monitoring was already available — eliminating redundant calendar-driven teardown inspections.
Approved recommendations were synced live to their CMMS (SAP PM), without the need for manual entry. The agent also flagged PMs tied to API/OSHA process safety requirements, locking them from optimization to protect compliance.

The Impact: From 1 Compressor Train to a Field-wide PM Strategy Optimization
After the Maintenance Optimization Agent identified ~$240K in potential savings on a single gas lift compressor train, the operator extended the program across remaining compressor trains and ESP systems field-wide, identifying cost savings of over $650K.

