Both OpenAI and Google have released major new AI models within days of each other. OpenAI launched GPT-5.6, and Google has followed up with Gemini 3.6 Flash. While both companies have talked up the coding and developer features of these models, they also power the consumer versions of ChatGPT and Gemini.
So, rather than measuring them with programming benchmarks, I wanted to find out which is actually better at the kind of messy, everyday problems most people use AI to solve.
For this comparison, I matched Google’s Gemini 3.6 Flash against GPT-5.6 Sol using its default Medium reasoning setting, since both are intended to be the standard high-quality models that paid subscribers will use for most tasks.
My digital life
So, I gave them my entire digital life for the week ahead.
I uploaded:
- Bank statements
- My calendar (they had access to my calendar app and I also sent in a screenshot of my calendar from another account I use).
- Grocery receipts
- Emails (both had access to my Gmail)
- WhatsApp screenshots
- Photos of my kitchen cupboards
- An energy bill
- Travel bookings
- Handwritten notes
Then I gave both models exactly the same instruction:
“Tell me everything I should do this week.”
It sounds like a simple request, but it forces an AI to combine information from multiple sources, prioritize what’s important, spot deadlines, reconcile conflicting information, and produce a practical action plan. In other words, it’s exactly the kind of real-world problem people increasingly expect AI assistants to solve.
The results were like night and day.
Two different approaches
Become a TechRadar Insider by simply clicking ‘Join Now’ at the top of this page. Have a question? Please email membership@techradar.com
When I uploaded the files to Gemini 3.6 Flash and asked what I should do this week, it largely ignored the photos and concentrated on the screenshot of my calendar. Its initial answer mostly repeated the events I already knew were happening. Thanks, Gemini — I had the calendar open in front of me.
I then had to explicitly ask whether it could infer anything useful from the other images. It eventually offered some additional advice and did a good job of dividing the information into categories, including work, shopping, fitness, notes, and receipts. But it failed to flag that I had two clashing events in my calendar that evening. It also offered very little prioritization or practical guidance about what I should do next.
ChatGPT took considerably longer to respond, but its answer was far more useful. From the WhatsApp screenshots, it correctly deduced that attendance at my Friday Tai Chi class was likely to be low and suggested I decide whether it was still worth running. It noticed that yoga had been canceled, and spotted the two conflicting events in my calendar, telling me that I needed to choose between them.
It also totalled the receipts I had uploaded, suggested what I should do with them, and made a decent attempt at deciphering my handwritten notes. More importantly, it organized everything into a day-by-day plan for the coming week, then identified the three most urgent tasks, so I knew exactly where to begin.
The crucial difference
That was the crucial difference. Gemini told me what was in my files. ChatGPT worked out what I should do with the information. In World Cup terms, ChatGPT scored a hat trick while Gemini missed a penalty.
Google says Gemini 3.6 Flash improves coding, knowledge work, and multimodal performance compared with its previous models. That may be true, but in this particular multimodal test, it was comfortably beaten.
When I asked both AIs to make sense of real life rather than pass a benchmark, ChatGPT reasoned about the information in a far better way than Gemini did..