Amazon Lookout for Equipment
Amazon Lookout for Equipment is a machine learning service that enables industrial customers to build custom ML models to detect abnormal equipment behavior using their existing sensor data — with an automated machine learning workflow that requires no data science knowledge.
Unlock the full case studyOverview
Industrial companies stream enormous volumes of sensor data into the cloud, but building ML models to catch equipment problems takes time, resources and expertise most teams don't have. Lookout for Equipment automates that workflow for developers with only basic ML understanding.
I joined in December 2019 as the sole UX designer — while also leading Amazon Monitron — and worked from defining the vision through MVP prioritization, customer PoCs and post-launch improvements, to the service's preview at re:Invent 2020 and general availability in April 2021.
Automating a Data Scientist's Workflow
The service compresses what would normally be a data-science project into a console workflow. A user points Lookout for Equipment at historical sensor readings in Amazon S3 — timestamps as rows, up to 300 sensors as columns — and the service ingests, aligns, fills gaps and de-duplicates the data in ten to twenty minutes, then grades every sensor as high, medium or low quality so problems in the data surface before they become problems in the model. Training needs at least three months of history and can finish in under fifteen minutes.
Almost everything past that first upload is optional, and keeping it that way was a deliberate line. A user can add labels marking known maintenance windows or past failures to sharpen accuracy, switch on off-time detection so the model isn't learning from hours when the machine was simply idle, and tune the sampling rate against how slowly the anomaly they care about develops. Each is a real lever for someone who knows their equipment — and each is skippable for someone running a first proof of concept. The default path had to work without any of them.
That automation is what made the design problem interesting. The people this service is for — plant managers, reliability managers, line leaders — are evaluating whether predictive maintenance is worth committing to at all. They are not going to read a confusion matrix. So the hard question was never "how do we show the model's output"; it was "what does this person need in order to decide, and what can we leave out?"
The answer shaped every evaluation screen: an executive summary of how many labeled events the model caught and the average forewarning time — how much notice you get before a failure — stated plainly enough to paste into a report. A timeline putting detected anomalies against known events so the two can be compared at a glance. And for any single event, the sensors that contributed most to the model's conclusion, ranked by percentage, rather than a raw data dump users told us they would hand to their maintenance team anyway.
Once a model is worth trusting, it moves to scheduled inference: new sensor data lands in S3 on a cadence, the model scores it, and the detected events are written back out to S3. That last step is the one we chose not to design around. Research kept saying the same thing — teams already run maintenance work through existing systems, and results are most useful flowing into those rather than into another dashboard to check. So viewing inference results ranked lowest in launch scope, and the console's job became getting a team to a model they believe in and then handing off cleanly.
Highlights
- Clarified a sprawling end-to-end journey into five main workflows, and drove launch-scope prioritization across user insights, business needs and technical constraints
- Ran 12 interview sessions and validated designs with 10+ internal users and customer calls before external access existed
- Designed model evaluation around four intents — easy to understand, flexible, consistent, accurate — turning raw science output into decisions users can act on
- Passed 4 AWS UXDR milestone reviews while solo-owning two services with the same launch date
Impact
One of the fastest-growing AI/ML services offered by AWS in the industrial field, with thousands of models deployed across large corporations — saving millions by preventing downtime. The full case study covers the process, prioritization frameworks and design examples in depth.
Want the full story?
The complete case study — problem framing, process, design artifacts and outcomes — is available with a password. Don't have it? Get in touch.
Unlock the full case study