The day I realised I trusted Oura more than the AI explaining Oura

I spent several months running two health products side by side. Oura on my finger, and Bevel drawing from the rest of my Apple Health data. I was not looking for prettier graphs. I wanted to know whether an AI health coach could genuinely help me make better decisions, or whether it was just a nicer way to look at numbers I already had.

If Bevel was ever going to work for anyone, it should have worked for me. I am already data driven, already tracking my recovery, already using AI every day for real work. I was happy to pay for a premium tool if it earned its place. And I did not want another dashboard, I have plenty of those. I wanted to be coached: told, in plain language, what the data meant and what to do about it. I walked in as the ideal customer, wanting to be convinced.

At first, I was. The concept is genuinely excellent, and I want to say that plainly before anything else, because this is not a hit piece. An AI that watches your recovery trends and talks to you about them like a person is exactly the sort of thing I have wanted for years. The pitch is close to perfect. The vision was never the problem.

The problem crept in slowly, and it was not the one I expected.

It wasn’t the numbers

I assumed that if an AI health coach let me down, it would be because it got the data wrong. A misread HRV, a sleep stage in the wrong bucket. Wrong numbers I can forgive. Every wearable is an estimate rather than a measurement, and I made my peace with that a long time ago.

What actually wore me down was different. The numbers were roughly fine. The AI stopped listening.

Not once, as a glitch. Repeatedly, as a pattern. I would ask a question and it would answer a different one, sometimes one from a conversation two days earlier. I would correct it, it would thank me, and then it would carry on exactly as before. Any one of these could have been dismissed as a bug. Together they formed a pattern, and the pattern is the whole story.

The evidence

The clearest example was a running argument about sleep.

One evening, after a genuinely rough night of four hours and forty-four minutes, and then a full day, Bevel told me my “Sleep Bank” was sitting at a deficit of minus seven hours and thirty-two minutes, and suggested a 9pm “screens off” alarm to start clawing it back. Oura, looking at the same nights, had my sleep balance far lower and nowhere near that dramatic. Same body, same data, two very different stories. This is exactly where a good coach earns its money, so I did what you are supposed to do. I told it, plainly, that the number was wrong and that Oura had me much lower.

What I wanted was for it to take the new information on board. What I got was a defence.

It explained, at length, why its figure must be right: a weighted sum of the last seven days, one very short night anchoring the total, and so on. Then it told me how exhausted I must be and reached into my personal life to account for the number. I had never said I was exhausted. When I pushed back and said so, it replied, “Fair. No assumptions, I’ll stick to the data,” and then produced the exact same minus seven hours and thirty-two minutes.

That is the part that matters. The mistake was never the problem, every system makes mistakes. The problem was that better information arrived and nothing changed. It defended the conclusion instead of updating it.

The corrections were the tax I paid for all of it. The same ones, again and again, with no sign any of them had stuck. Then tried to tell me that I was exhausted, this was the last straw, I stopped trying to fix the product and started trying to leave it. I was not angry. I had simply run out of reasons to trust it, and once that goes, a coach is just a confident stranger talking over you.

Trust Comes From Attention

The reason I trusted the ring more than the coach had nothing to do with which one was smarter. Bevel is, on paper, the more sophisticated product. It is the one doing the talking, the reasoning, the interpreting.

Trust does not come from intelligence. It comes from attention.

There is a difference between performing understanding and actually demonstrating it, and once you have had both from a machine you notice it everywhere. Performing understanding is fluent, confident, and pointed at the wrong question. Demonstrating understanding is duller and far more valuable: it tracks what you said, holds the thread, and answers what you actually asked. Bevel could do the first all day. It kept failing at the second, and the second is the entire job.

For what it is worth, I ran some of the same questions past ChatGPT, which knows nothing about my recovery and holds no special health expertise, but at least remembered what I had just told it. The generalist held the thread better than the specialist. That should worry the specialist.

Oura, by contrast, mostly shows me the data and gets out of the way. It does not perform. It does not reassure me. It hands me a number and a trend and trusts me to be an adult about it. For a long time I read that restraint as a lack of ambition. I have come around to seeing it as the more honest design. The less the wearable tried to interpret for me, the more I trusted it, and that is an uncomfortable lesson for anyone building conversational AI.

I want to be fair, because the idea is worth rooting for. I genuinely hope Bevel succeeds, because the concept is outstanding and the market needs someone to get it right. If the attention problem gets solved, I would happily pay again and revisit the whole thing. This is not a company I want to see fail. It is a product I could not yet trust.

What I actually took away

I do not think Bevel’s model lacks intelligence. I think it lacks sustained attention to the conversation, and I have started to suspect that is the real frontier, not just for one health app but for most of what we are building. We keep reaching for bigger models, as if the next jump in raw capability is what will finally make these things trustworthy. My few months with an AI coach point the other way. The gap I felt was not a knowledge gap. It was an attention gap.

An assistant that listens carefully to a small context will out-trust a genius that keeps losing the thread. Every time.

Which leaves me with the question I cannot shake: as these tools get smarter, are we quietly measuring the wrong thing?

Leave a Reply

Discover more from Paul Usher

Subscribe now to keep reading and get access to the full archive.

Continue reading