Published Sep 19, 2026, 2:00 PM EDT Mahnoor Faisal is a tech journalist covering AI and productivity tools with bylines at XDA, SlashGear, MakeUseOf, Laptop Mag, and Android Police. She's been writing professionally since she was sixteen, and has since penned hundreds of articles. This includes in-depth coverage of AI tools like NotebookLM to breaking news across the AI space. Her passion for technology started when she received her first iPod Touch (4th generation) on her 8th birthday, and she's been deep in the tech world ever since. Currently pursuing a degree in computer science, Mahnoor brings both a journalist's eye and a technical foundation to her coverage of how AI is reshaping the way we work and learn. I've tested countless AI tools and sat through what feels like a million briefings at this point, and something I keep coming back to is just how few AI use cases seem designed for the average person. We're constantly shown agents that can string together elaborate workflows, models that can write entire codebases, and increasingly specific features that sound impressive in a demo but aren't necessarily things most people will ever need. One use case I've never felt that way about, though, is documents. Give an AI a ridiculously lengthy PDF, ask it a handful of questions, and have it pull out exactly what you need. This is exactly the kind of thing I can imagine practically anyone finding useful. NotebookLM has been my favorite tool for this use case, but ChatGPT and Claude have become much better at working with long documents too. So, I decided to put them to the test by giving ChatGPT, Claude, and NotebookLM the same 100+ page document. I tested all three with the same document and questions Same PDF, same questions, no excuses I hunted down a 113-page-long document about a railway collision in England that went into everything from the events leading up to the crash to the technical and organizational factors behind it. The document also had a good amount of visuals and graphs in it, and there were also some solid opportunities to test whether the tools could connect information scattered across different parts of the report. So, I fed the document to all three tools. To keep the test as fair as possible, I used the premium versions of all three. These are the questions I used: What time did the collision happen and how fast were both trains traveling? Explain what happened in five sentences to someone with no railway knowledge. Reconstruct the events leading up to the collision in chronological order. What was the immediate cause of the accident, and what factors made it possible? Which pieces of evidence support the report's conclusion about why the train failed to stop? Who governed the design of the TPWS installation? Based only on the report, what intervention would most directly have prevented the collision? Explain using evidence. Look at Figure 41. At the date of the accident, was the Ellipse work bank above or below the red dotted mean-average line, and approximately how many tasks separated the two? All of these questions are meant to test a slightly different part of document understanding. Some are straightforward retrieval questions, while others require the tool to summarize, reconstruct a timeline, connect evidence from different sections, reason from the report's findings, or accurately interpret a chart. ChatGPT ended up being the surprise winner ChatGPT pulled ahead when it mattered I've spent a lot of time pointing how ChatGPT has fallen behind competitors in the last few months, but recent times have proved me wrong. This experiment was yet another example of how that's no longer the case. For starters, out of all three tools, ChatGPT was the only one to give me the precise time of the collision: 6:42:57 PM, rather than simply rounding it to around 6:43 PM. It’s a tiny detail on its own, but it was an early sign that ChatGPT was digging a little deeper into the report instead of stopping at the most obvious summary-level answer. The advantage became much more obvious as the questions got harder. When I asked it to reconstruct the events leading up to the collision, ChatGPT was the one that went most in-depth. It was even better when I moved away from simple retrieval and started asking questions that required actual synthesis. For example, when I asked what evidence supported the report's conclusion about why the train failed to stop, ChatGPT pulled together the train's own recorder data, the wheel-slide protection activity, physical testing of the contaminated rails, post-accident braking tests, and the investigation's findings about where the driver began braking. Instead of treating one paragraph as the answer, it connected evidence scattered across the report to explain why the investigators reached their conclusion. Claude’s biggest mistake sounded completely convincing This is where Claude lost me I intentionally included one question from around the 80th page of the PDF just to see whether the tools were actually searching through the entire document, rather than relying heavily on the summary and the sections closest to the beginning. All three tools managed to answer it successfully, but the biggest difference showed up when I asked what intervention would have most directly prevented the collision. That question required a lot more than simply finding a relevant paragraph and repeating it back. The report actually models different braking scenarios, and ChatGPT picked up on the most useful one: RAIB found that, given the conditions that night, the accident would probably have been avoided if the driver had started braking at Broken Cross bridge. It also distinguished that from braking slightly later at the fallen tree, where the report was much less certain that the collision would have been avoided. That kind of nuance is exactly what I was hoping these questions would expose. Interestingly, Claude straight-up got this one wrong. It picked double variable-rate sanding (DVRS) as the most direct intervention and even claimed the report didn't quantify what would have happened if the driver had braked earlier at Broken Cross bridge. Except it did: RAIB explicitly modeled that exact scenario and concluded that the accident would probably have been avoided if braking had started there. That made this one of the clearest differences between the tools for me. Claude’s answer sounded convincing and was backed by relevant evidence, but it missed a more direct piece of evidence elsewhere in the report! Another aspect of Claude's responses that disappointed me was how complicated some of its answers were. For instance, even when I explicitly asked it to explain the accident in five sentences to someone with no railway knowledge, it crammed in more technical terminology than needed. What frustrated me here was that the response was almost poetic and overly polished rather than actually simple. It sounded nice, but that wasn't really what I had asked for. I wanted the kind of explanation you could give to someone with zero railway knowledge without making them stop and decode half the terminology. NotebookLM’s citation system is still hard to beat Credit where credit is due ChatGPT might have won me over in terms of the depth and quality of its answers, but NotebookLM still has the edge when it comes to citations. All the responses you get come bundled with clickable citations, and hovering over them shows you the exact passage the answer was pulled from without forcing you to go hunting through the document yourself. While you can technically ask ChatGPT to point you to where it picked up the information from, NotebookLM's citations have spoiled me! It's worth mentioning here that NotebookLM didn't exactly disappoint me with the quality of its answers. Unlike Claude, none of its answers were factually incorrect or overly wordy. However, relative to ChatGPT, they often felt a little more surface-level. It usually found the right information and explained it clearly, but ChatGPT was better at pulling together details from different parts of the report and adding the kind of nuance that made its answers feel more complete. Ultimately, the results of this experiment were more interesting than I anticipated! While I'd still lean on NotebookLM for situations where I really don't want hallucinations, I think ChatGPT is giving it some serious competition when it comes to PDF analysis.
I gave ChatGPT, Claude, and NotebookLM the same 100-page PDF — here's who actually understood it
Full Article
Original Source
Read the full article at Xda-developers →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.