I thought LLMs were a great tool for learning new topics - perhaps even complex ones. But overtime, I have had several frustrations with this. First, I get exhausted reading LLM prose. I really don't want to read anything generated by something like Opus 5 at this point. Second, as I dive deeper, I need a way to organize the information in a useful way as I begin to branch out in many different directions. I have tried to use the LLM to fix this by having it generate a web page with diagrams and organized information flow. It's an improvement, but I still run into the issues I described in my first pint - LLM prose is annoyingly dense, and the useful information gets lost in a bunch of noise. You can direct it do something like "use plain English and avoid LLM prose - provide only as much information as necessary to demonstrate the point", but it is once again only a marginal improvement.
And then I begin to think to myself that I should just read a book on the topic written by a trusted source who put a lot of effort into teaching the topic properly and presenting the information in a thoughtful way. So, I am back to books and mostly try to use LLMs to clarify certain questions or ideas I have.
LLMs can't "read the room" and infer how much context the audience already has, so they try include everything.
human conceptual thinking is very much a multi-dimensional graph, which relies on light "approximate" concepts that are "good enough". LLM AR token generation is extremely one dimensional and doesnt care about the "weight" of the concept behind a token.
LLMs hold billions of parameters in "mind" at once. humans hold like four "concepts".
This is the essential mismatch and the primary reason LLM conversation can be so painful and exhausting.
Explaining this and limiting "concepts" to four at a time tops is one of the very few AGENTS.md / system prompts I always use, and it has proven invaluable time and again.
Thinking traces show how effective this is at forcing the LLM to simplify its thinking.
[edit] Also, myself and nearly all of my peers are struggling to choke down the flaws of LLM tooling along with the benefits. the speed at which LLM adoption is being forced, without truly crafting them into quality tools first, is not ok, and not normal.
LLMs have stirred an inhumane hunger and fear. the tech is fine, but the way tech companies (creators and consumers) are behaving should be deeply questioned.
If "attention is all you need" then it's something we do indeed lack, in comparison to LLMs! But it's an interesting question: might machine cognition benefit from similar bottlenecks in an attention algorithm? Advancements like Kimi Linear seem to indicate that we're far from the finish line: https://arxiv.org/abs/2510.26692
It's much better to feed the book to the LLM and ask questions as you read along, instead of asking the LLM to basically write a custom book for you from scratch.
I have to agree with this. Completely relying on LLM for all your learning needs is a disaster. But being absolutely against use of LLMs isn't doing you any favors. This is where you don't have a formula but rely on you judgement and evidence of your having learnt something.
For example having an LLM summarize a dense topic and to find books so that you can filter faster and spend time reading those books works way better than having the LLM summarize the books or the topic (or even relying on second hand information). Another one is having the LLM quiz you on your topics of interest. With questions tailored to attack specific areas that you struggle with. Its wonderful at this, nothing I've used comes close to what an LLM can do here.
You define for yourself what your goals are, slowly refining them as you learn more, and use LLM as a tool. This ,I find works best for learning.
Make it your goal to teach a room full of other humans that topic. I guarantee you will know that material cold. I've done lots of technical training in my career and after teaching a class two or three times I find myself to be very competent in the topic.
It's long been the case that the best way to learn something is to teach something.
> It's long been the case that the best way to learn something is to teach something
Which is pretty unfortunate for those that want to learn. I used to enjoy writing documentation at work, it was my favorite part of the job. And it did feel like it benefited me more than it benefited all the people that were (or weren't) reading my documentation. Now I can't really justify spending much time on docmentation when LLM's can do it in a fraction of the time and it's "good enough"
I got started as a software engineer working in the nuclear industry in the 80s. We measured our documentation in inches not pages, and it was all written by hand. And I'll bet you the documentation in the nuclear industry is still written by hand and not by LLMs.
There is still value in experts distilling knowledge and crafting it to the audience. I'm a consultant in cybersecurity and someone asked me "give me a best practice framework for good policy hygiene". I'm sure an LLM could spit out some tips, but I've been on the industry 15 years and can write in 5 pages what an LLM wouldn't conceive of in that space.
Even with the latest models today, the hallucination rate is absurdly high on anything deeper than surface level knowledge or something that can be directly scraped from reddit.
And you notice when it's a topic you know well or something like software where you can immediately tell the options it's giving you don't exist on the page. Leading to the amusing statement "LLMs are bad at what I do but great at everything else".
I've (elsewhere) written about this diminishing return effect on LLM utility in relation to increasing expertise.
The question: what's the net positive gain of turning people who know nothing in a given field into sub-novices, while weighing actual experts down with work slop and marginal returns?
And I wonder what the true cost is of arming so many novices with that level of dangerous knowledge.
The fact that some people get genuine value from LLMs when learning doesn’t contradict the fact that they’re Dunning-Krueger “expertise” generators. The fact that the person learning from them is in charge of ensuring they aren’t full of shit, which they frequently are, is an inescapable flaw in this process. I honestly think that reduces the value of these things to just above what you can find out with a search engine with most topics. Hey, great. An improvement is an improvement right? Is it an improvement worth trillions of dollars and screwing over writers and artists worldwide? Fuck no.
It wasn’t even all that long ago. I noticed around the time ChatGPT 3.5 was released I saw a significant change in tone regarding LLMs and diffusion models. I think that’s when some serious astroturfing started.
I’m old enough that I worked my first IT summer job the same year slashdot was founded. I’ve seen a lot of tech tribalism form and dissipate, and this one didn’t feel organic. My gut says a lot of the us-vs-them tension originated in a deliberate campaign to cast AI boosters as the tech industry in-crowd, and ‘other’ the people not on-board. Who knows.
I came across the socratic method recently, and have used it to learn a couple of topics that I was having trouble getting to stick. There are some SKILL.md's available for it. It works for concepts as opposed to facts, and causes the model to guide you to answers through your own reasoning, which is both much more engaging than reading a wall of LLM text and helps the information stick.
The "Socratic Method" (aka maieutic) skills annoy me, precisely because when you read them they are the kind of low-effort, low-expertise crap someone who over relies on AI would naively come up with when tasked with the problem of coming up with skills for learning. "Hey the Platonic dialogues are pretty cool and smart, let's do that".
The body of literature on learning theory, and beyond that on specific types of learning and specific mediums such as learning from text is so rich there are way more useful models to draw from. Believe it or not, prellm, researchers in the textual learning field had already demonstrated you can achieve performance equal or better than novice tutors using pretty basic computer aids that follow specific hint/pump interaction structures. Guiding an LLM to use these findings has evidence backing it and is way better than telling it "i guess be like socrates". The problem is, to realize there might be richer more effective and highly researched ways of tackling the problem beyond the first fart of a thought you had one afternoon requires the deep respect for expertise and specialization that precisely basically everyone in the AI space right now fundamentally lacks.
what a time to be alive! "if you're having trouble understanding what your robot tutor is trying to teach you, you can ask it to guide you to the concepts using your own reasoning. This is both much more engaging than reading a wall of the robot's text and helps the information stick."
I also agree with this. LLMs are a great companion when reading a book to clarify things and dive into specific topics.
I'd imagine an application that uses LLMs will be created that better manages learning. It's just not clear what that UX is yet- it's obviously not just a chatbot
Whats funny about LLMs is they are trained from books, but they are also trained to not output books, so how much of an LLM skews its output because a perfectly normal sentence could be a quote in like 300 different books?
i threw the entire sanderson cosmere into a RAG graph sorta deal just to see how it would do if i questioned an mcp server for it about a universe i know decently well. it was actually astoundingly good. was able to find easter eggs acrossed different books and answer dumb questions like "why is kaladin emo"
If you don't know why kaladin is emo, did you truly read the books lol. Every character has to deal with the stresses of war and most don't come equipped with good mental health to begin with, they're just normal people
You need a non DRMd copy of the book. You don't have to feed it all at once, although with a 1M context limit it is doable. A few chapters at a time is enough, in my experience. An easy alternative is using NotebookLM (now Gemini Notebook), and that has worked brilliantly for me, but I haven't tested it for technical topics (for that I like the LLM to create graphs and e.g. interact with Mathematica, so I haven't tried it).
LLM is still too verbose, a real person Socratic conversation can interact a couple sentences at a time, not spew 1-3 windowfuls of low density bullet points.
I even wonder if this behavior is due to next-token prediction architectures, somehow.
This. This is exactly how I use them and I have had no issues so far. I read the book myself, then I point the LLM at it to ask questions about notions I might be struggling with.
I agree, I too like to read authored books. But there are certain topics, especially the new ones doesn't have good books yet. My post was to show that you can create books based on your interest, in a way that you like w.r.t to the content/tone/layout etc. It's a new idea made possible by LLM's. It might become the norm in few years I believe, where each one will be having a personlized library of books that they curate. So the point was to persuade the parent commentator that some books are better created this way.
>I really don't want to read anything generated by something like Opus 5 at this point.
Personally, I find that its generated prose tends to have an undue weight to it, almost as if every topic I ask about somehow bears a heavy burden, or is otherwise load-bearing, to use its parlance.
I have a personal theory: LLMs are *fundamentally* handicapped at perceiving what's going on in the mind of the human (this can't be "innovated away") and that's at the root of what makes them suck at conversation.
Next time you're chatting with someone, notice how much understanding is shared without anything being said. E.g. the other person might share something deeply disappointing, and they can tell without you even saying anything whether you get what they're going through. This unspoken-yet-communicated information guides the conversation. Or as another example: humans can read the room -- you walk into a room and immediately adjust your demeanor based on what you see and sense.
LLMs are totally blind to things like this, and this adds an inescapable awkwardness to interacting with them. I don't believe they'll ever grow out of this. Which thankfully implies more long term demand for humans instead of robots. :)
Very well observed. I found one more thing: they fail to consider what a 3rd person might understand from your conversation, so when you ask it to dump stuff into a Documentation, they keep making references to facts you had previously discussed or to the train of thought, completely irrelevant to bystander.
100% -- this is the worst. Referencing all sorts of "words with made up contextual/analogous meanings" based on the conversation...outside of the conversation.
Does anyone have a read on if this is primarily a Claude issue, or if all LLMs do this?
It has a sort of metronomic quality. It never slows down or speeds up or modulates its tone. It plods forward at a relentless pace and never has a light touch with anything.
I think this is one reason why LLM text is pretty exhausting to read for long stretches.
>It has a sort of metronomic quality. It never slows down or speeds up or modulates its tone.
It's possible that this quality you describe stems from the extensive training corpora utilized by the major AI labs. These almost certainly include work from the esteemed economist Jacob Silj:
I never used office hours as a student, which I later regretted because it made me work longer and harder to perhaps achieve somewhat better understanding in some classes, but also I dropped every proof-based math course I ever took. Overall I think my education would have been stronger by attending office hours.
I view LLMs in education similarly to office hours. Some people abuse it to get homework answers without grappling with the material, but the optimal amount is not zero.
LLM certainly not a replacement for a book, where you get someone’s extended personal approach to a topic, thoughtfully organized, reviewed and edited, often times actual courses taught based on it, with answers checked and errata available online.
I completely agree. I have vibe coded what i would consider to be some pretty weird things in the name of learning facilitation.
Perhaps the best example has been a native macOS app that is a completely custom text editor with built-in debugger, lsp support, fuzzy finder, etc stuff you'd expect. Inside the same app is a library of books i can read within the app completely formatted and for every chapter/section of each book that is a quiz to take (LLM generated of course), a "recitation" tab where i am asked a question and say outloud my response to the AI to evaluate me on and then finally practice problems to do within the custom text editor (these are usually programming books). The reader also has ai re-write built in.
As neat as this is, and i worked through K&R like this, i have ultimately fallen back on "just read the damn book and go to the AI when you've got questions."
I have had very similar experience! I wanted to learn Probablistic ML, checked out a couple of MOOCs, but didn't find any that were at my level - some were too advanced, some too beginner level. Claude was unable to one-shot a course, so I am now asking it to generate it module by module. But even here, it is not doing a very good job. I muddle through the concepts that it has written, do a whole bunch of back-and-forth, which tbh is exhausting, and then rewrite everything in my words so it actually makes sense to another human being.
> I get exhausted reading LLM prose
So much this! If I see one more sentence with the words "genuinely" juxtaposed with "load bearing" my head is going to explode!
i’m personally deriving a huge amount of value from the custom materials fable is assembling for me. for example i asked it to write a focused expository math paper on reed solomon to accompany an implementation module that it wrote for me. it’s remarkably useful to steer it to create graphs and diagrams of exactly how you like the material presented. or the bibliography researched and cross-linked with the body or the order you want your questions addressed.
it also researched vision correcting displays for me and i can finally put that idea to bed - i was never really going to pick up an optometry textbook tbh. plus it was able to pull together a bunch of geometric and physical context about light and the eye plugging exactly my personal knowledge gaps.
in general i suspect these materials might not be that interesting to others because they are so custom to my learning style and personal needs and preferences.
these are usually not one shot documents but rather many prompts deep before i get something I’m willing to sit down and read or study. but dramatically quicker than assembling it myself from primary sources. i wouldn’t say it matches master expositors but then they’re not available to write on any topic i happen to need right now.
plus I’ll just have a live voice discussion with the system when i go for a walk and there are still things bothering me on a topic. it takes a little patience but if i’m in the mood it’s amazing.
i generally find that it can help track down specific references if i suspect hallucinations. but especially on factual topics my experience so far has been extremely encouraging.
Good point. Same with me. Despite it being very tiresome, all the back and forth that I do with it really deepens my understanding of the topic. I have been on Opus so far, let me try Fable and see if it gets better. I haven’t tried voice either. Next time I go for a walk I’ll try that!
Good to know other people who get migrane reading LLMs dense prose. I started reading books again recently, since everything online is polluted by LLM prose. What i realise, is that a human author, especially a teacher understands the learning pathways of new learners, they motivate the learning, and start from simplest concepts (a spherical cow), and then building all the complexities. This helps us to emphasize on most important concepts, while throwing away unnecessary complexities. While reading LLM prose is like reading a research article, that is written to an expert in the area, that talks about bleeding edge, with full of jargons, caveats, that just is not conducive to the learning process for a new learner.
Also they tend to assemble complex jargon in obtuse or meaningless ways, which makes reading and parsing and understanding much more difficult. Tends to reveal that LLMs fundamentally do not have "understanding", just likely word generation
To me, it's just Claude. The other models have their quirks but nothing is quite like Claude.
But even with Claude, it's it's really the prose getting in the way you can install the caveman plugin or tell it to use that "standard technical English" thing.
I wish I had something more methodical I could show. It's all subjective, but GLM-5.2 feels more human to me. Even GPT-5.6 Sol tends to be easier on the eyes for me (though the stereotype of it overengineering and no common sense are still true).
I tried using a new agent service recently and could tell immediately that it's powered by Claude due to the way it writes.
I've decided that the main thing I'm building is my own mental model. You can take notes, create docs, put graphs and websites together, but unless I'm just trying to generate some reference material the only real objective is to develop the understanding and intuitions inside my own brain.
So I have the LLM offer a very short explanation of something, and from there's it's just me asking questions. Anything that feels fuzzy or not fully internalized is something I poke at until I'm satisfied.
It really has helped me develop a sensitivity to what I understand vs what I don't, and the ability to drill into any part of it is amazing.
I do this too. I use Claude. I picked a voice I like. I go on a three mile walk. I will ask it questions about a topic that I want to learn about. If it starts telling me more than I want to hear right then, I will say "stop". It doesn't get offended. I then ask it something else. I find this very effective. I control it so it only explains to me what I want explained. If what it says sparks questions on a related topic I jump to a brand new topic. No personal tutor could keep up with this or adjust to exactly how I want to be addressed like Claude does. I'm very excited about the progress I'm making mastering new topics.
And yes, it is not that it is just presenting the facts. By me taking control of the direction the questions and answers go, I can flesh out my mental model. I won't retain every little thing it tells me. But I am much farther ahead than before.
Claude code has a „fork“ feature where you can fork an existing conversation and keep talking in the fork and then you can go back to the original of the fork. You can fork as many times as you want and let LLMs write to a markdown file to keep important facts and learnings - also good for agents to do research without expanding the context window
> First, I get exhausted reading LLM prose. I really don't want to read anything generated by something like Opus 5 at this point.
Agreed.
I find Opus 5, and even Fable, to be overly wordy in eg PR descriptions and code comments.
However, I suspect that's more to do with what they are trained to do by default than LLMs in general. I have a little setup where I tell Claude to work together with Codex to tighten up prose and comments, and for me that produces much more palatable text that needs less human editing afterwards.
The sycophancy is also a concern, it’s not really an impartial teacher, all its training is to suck up and maximize engagement rather than learning. The incentives are wrong.
I have been using LLMs to help me turn my journals into interconnected notes and sometimes it is so confusing to read the notes that it doesn't resemble any human would write. Its like the models are getting stronger while also losing its touch to write human sounding sentences on complex topics.
I think one problem is that books are not customizable, and many books are aimed at people with some certain knowledge. With LLMs, you can tell it what your knowledge level is and ask it to customize the answer for you. This is difficult to achieve with books.
While my advice is specific to learning about codebases, the way I do it is to have it generate mock data and put it in the local development environment, and give me some exploratory commands, and then ask away. It's a machine after all, so I don't have to read its preceding prose to understand whether it did tell me something, it can just repeat it however many times I ask it, and the hands on commands etc. give me something to actually try and implement.
I use the LLM to point me at relevant books and papers. But there is still a trust problem: I am trusting the LLM to point me at reliable, trustworthy sources.
I'm not sure I'm better off with humans though -- I'm not qualified to judge whether a source is a proper authority, not an I qualified to judge whether someone knows enough to point me to a reliable source.
It seems this is a fundamental epistemological problem to which there may never be an answer.
> LLM prose is annoyingly dense, and the useful information gets lost in a bunch of noise
This is my biggest gripe with reading AI-generated text as well (ignoring the meta issue of whether it's worth taking the time to read something that an author didn't think was worth the time to write). It's gotten to the point that weird AI-style analogies just take me completely out of the text and kill my interest.
You can make something like the link below. I'd say it's a more than marginal improvement. It's still tiring, but it's much better according to my taste. I imagine everyone would have their own version of this for their own preferences.
It's a loop that uses adversarial review to check several dimensions of the writing:
Even if you tell the LLM not to use LLM pros they (still) do it. If you feed the Wikipedia article on signs of AI writing and tell them to use none of those signs they will also (still) do it. I have tried (many times) to get an LLM to explain a concept to me, or a process, or an algorithm or what have you, and every time they cannot help themselves. Either they use LLM pros, or they get so verbose that it all just becomes noise and I spend more time filtering out unnecessary jargon than I do reading let alone learning anything.
I've found it helpfull to ask the LLM to generate a sylybus for the topic. treat the sylybus as a design doc for a price of software. i find they do much better when they have subtasks to focus on. they can do big picture and small picture, but they can't do both at the same time.
> First, I get exhausted reading LLM prose. I really don't want to read anything generated by something like Opus 5 at this point.
I’ve found the tone of Kimi K3 to be less obnoxious. Unfortunately it doesn’t wholly solve the issue, I don’t think any LLMs out there have a truly pleasant writing style, but at least not every assumption is “load bearing”.
Same. I usually just read the book along and ask questions on a specific part, rather than trying to get the LLM to produce an entire study guide for me. It seems to work better as a Q&A than a "teach me" advisor
Yeah, I am currently trying to work with Opus 5 to refresh myself on deep learning fundamentals, and... it's a mixed bag. I'm glad I already am familiar with the subject matter, as I can prompt for refinement and improvement. It is kinda following the Karpathy videos so far (a couple lessons in) but adding more math/derivations, which was what I asked for. It has trouble staying on topic, presenting information in a coherent/meaningful order, and providing all the context necessary to move through steps in its "course notes".
Like I said, I'm essentially continually prompting to refine the material. LLMs certainly continue to append, and never cut back. It just keeps spitting out additional content at me. So that's a bit annoying too. But I can basically get figure out what's going on with a few extra promps.
If youre curious what i've got so far... just be warned it is quite literally AI slop plus me continually prompting for clarification/cleanup etc. : https://github.com/cmoscardi/ai-for-ai
opus 5 doesn't follow instructions that great. probably needed that extra "creativity" to benchmax. if you are stuck on anthropic, try Opus 4.8 or Fable5 a try with the same prompts. very different results.
To get rid of the llm prose issue you can take a representative sample of its prose (say a question and its answer), rewrite the answer in the way you'd prefer, and add that to the system prompt; I did this to get my LLMs to compress down what they say, and now everything they say is very dense and to the point
anybody adds some prologue about what style of answers are prefered ? i know i often try to change the linguistic patterns because i too (unsurprisingly) am tired of llm prose.
>LLM prose is annoyingly dense, and the useful information gets lost in a bunch of noise.
This problem doesn't get talked about enough and is second only to the hallucination problem IMO.
AI produces so much noise to wade through in order to find signal, and the more expertise you have in a field the more that costs. That noise directly subtracts signifcantly from productivity gains.
And, I think the problem is directly related to the hallucination problem. It feels very much like an effort to kitchen sink the response in order to provide some value among possible hallucinations.
It also seems to be a byproduct of Gen AI operation. It just fundamentally doesn't understand what it's outputting, so doesn't know how to narrow down to the most salient bits.
something to try that works really well is to put "explain this as if you're talking to a 5th grader" at the end of your request.... it just breaks down the text into more manageable sentences that can be understood by general audiences
It's just long. It just doesn't shut up. It's overly verbose. And you can't tell it to be concise or you degrade its quality.
If I ask what an integral is, the correct answer is that it is the continuos analog of a sum, generally used to calculate areas and volumes.
It should really be a single sentence, and then let me ask more about the terms I don't understand, and here's the beauty, in the previous one there can be only 5 terms I cannot know.
An LLM will vomit an entire page or more of explanation which isn't bad per se, but is an answer to something different: "give me a short introductory explanation to integrals". And that's not what I asked.
I call that vomit shotgun answers, text from which you have to filter out all the extra info the LLM wasn’t asked for. Luckily you can control that behavior and make it behave closer to what you want. Just ask the LLM how to ask for it.
The new Google Translate. They've made it slightly better at translating paragraphs of text but in many cases it's lost the basic function for translation: dictionary.
Try it out, fairly sure that if you out in 100 random words for 30 of them it will just refuse to translate them (it will copy paste the original word into the target language) or it will do silly things like use the target 4th dictionary definition instead of the primary one).
I think AI is surprisingly good at this. I use voice mode while working out to learn complex topics, follow up with reading, and then go back to ask the LLM more questions. They excel at simplifying complex ideas and have endless patience. One hack I found is telling the LLM to test my knowledge by asking me questions—that gives me a clear idea of what to read next. Overall, they’re a great tool to use alongside traditional learning methods like reading books and working through practice problems.
This is on top of the issue of LLMs being a moron. like, I'm sure dumb people can learn stuff from them, but I'm sticking to books written by people who know what they're talking ahout and not vibed slop.
AI hater here. Too me this story is all to predictable. This is exactly how I thought using LLMs to learn would go. Or in simpler terms: well duh.
I'm actually going to make a prediction here as well. I think you will soon realize that using LLMs to clarify certain questions or ideas you have will turn out to have frustrations as well. And that you will soon direct those questions to either peers you know in real life or internet forums which are very likely to have a non-AI policy.
> I'm actually going to make a prediction here as well. I think you will soon realize that using LLMs to clarify certain questions or ideas you have will turn out to have frustrations as well.
Much more likely that people will believe themselves to be an expert in a subject after having had a conversation with Claude about it.
Wild because people wouldn't do that with a friend. Something about an LLM giving people info has this insidious stoleb valor to the LLMs work. Like that wasn't your thought, you just google searched it.
Specifically, you can set the tone and style to match what you like or used to after asking it to interview you, create an output, and then use that output as the project description or document to refer to
If the output can't be trusted, and you use another llm whose output can't be trusted to check the untrusted output of the first llm, then you're back where you started.
Yeah this seems to me similar to how the mortgage backed security risk concentration occurred leading up to the global financial crisis. Whereby the risk from exposure to low grade / risky single mortgages was eliminated via diversification but the diversification was simply packaging all of the risky MBS’s together and in no way diversified or de-risked the entire portfolio
I'm hoping that the Big Horrible Realization comes sooner rather than later, when we have less collective damage and pain riding on it. (Plus I'd feel personally vindicated.)
But the problem with LLMs is that they get the facts wrong. PhDs are PhDs because they’d look it up in an authoritative source, or actually find out through research and experimentation. The whole point is that facts aren’t a matter of opinion. The only people that argue over documented, findable facts are idiots that nobody should listen to.
Not really. Take hallucinations for example. If they are 1 in 100 (actually they are much rarer, but for the sake of argument), then the chances that 2 LLMs or even just 2 runs of the same LLM have the same hallucination is, well, a lot less than 1 in 100.
Are there any reproducible hallucinations on any of the currently available OAI/Anthropic models? I’m not aware of any.
And even if they are related - if Opus 4.8 always has a 1:100 chance of a specific hallucination - then running the same model twice does indeed dramatically reduce the odds of an error in the final output.
If simply running things thrice-over was enough to stop "hallucinations" (and not incur other problems) we wouldn't be here talking about it today, it'd have been "solved" months or years ago.
"Turtles all the way down" is a phrase of, I think, unknown origin (https://en.wikipedia.org/wiki/Turtles_all_the_way_down) about infinite regress or trying to patch up some bad theory by appealing to itself. Someone claims
that what holds the Earth in place is that it sits atop a giant turtle, and a skeptic asks what holds the turtle up, and the response is that it's turtles all the way down.
Personally I think this is a bad characterization of using LLMs to fix up LLMs because while you can never guarantee results this way (as the quoted line claims here, which is worthy of criticism), it is, in practice, useful to use LLMs on top of LLMs. And there's no infinite regress. Auto-mode in Claude Code, for example, seems to me like it's been successful at making the system more safe than --dangerously-bypass-permissions without prompting the user for permissions constantly.
There are certainly uses where it’s good enough, but you can never be 100% certain of correctness in the way that people claim you can by stacking N layers of these models.
What triggered my response was the “just review the output with another LLM and it’s perfectly correct”
I read it as “if your LLM is being checked by another LLM, well then you need another LLM to check the checker. And can you really trust _that_ LLM? Probably should have an LLM to check the third one, and…”
It’s a reference to Bernard Shaw, who once said that if we ever created a truly artificial mind it would be inside a turtle’s shell. Sturgill Simpson covered the track on his seminal work, Xeno’s Paradox.
This guy remembers what I thought I was saying. In my defense I stole the whole thing from Stephen King’s It which I read … forty years ago, that can’t be accurate. Let me sort out my instruments and get back to you.
Agreed. Given how many significant errors LLMs make in my topic of expertise, despite my taking multiple error checking steps, the idea of catching 100% of hallucinations because you told the LLM to check itself is hilarious. It’s just a wild lack of insight: “I’m using the LLM to teach me something I don’t know about, I definitely have the knowledge base to spot any errors that might remain!”
Yeah. People with technical and/or tech business bonafides claiming that AI granted them expertise are so often taken at face value when they really shouldn’t be. Who told them that they were proficient — a chatbot? Someone who knows even less about the topic, so any expertise seems impressive? I’ll bet it wasn’t someone that actually knew what they were talking about. Even some tech reporters are tripping over themselves to be amazed, but don’t bother checking if they should be.
People don’t even have to be lying to be wrong about this stuff. Someone can learn enough about a topic to be halfway up Mt. Stupid in no time flat, and in doing so, think they not only truly understand the topic at hand, but might be particularly adept because they were such quick studies. People that know less are impressed, because why wouldn’t they be? Anybody that knows more than them sounds like an expert. And people that know what they’re talking about cringe at the overconfidence, and probably try not to engage: who wants to have to prove that someone’s boundless confidence is entirely baseless? Most of the time, they think the actual expert is full of shit because they think they’re the expert. It’s incredible how many times I’ve had people in tech confidently, even smugly “explain” design concepts and strategies to me that they did not actually understand, knowing I was an experienced, degree-holding designer… and they didn’t even have a chatbot’s lips on their ass telling them how smart and insightful they were.
Completely agree. While I didn't set things up to have AI review its output in a loop, my experience trying to build a specific acoustic testing rig with Opus 5 also aligns with the other "it's turtles all the way down" comment.
Opus 5 first built me a detailed plan, but a couple important details were either obviously wrong or felt unnecessary. I went back and forth asking for sources and more information probably like 4 times and every time it did the "in looking at things in more detail it appears my previous advice was incorrect" spiel. It just became exhausting at some point because it feels like it really lays bare how LLMs are just minimizing that loss function but don't actually "understand" anything. It was really useful as a search engine (it correlated some highly relevant source docs), but I just couldn't trust it to believe it was actually done at any step.
A second agent reviewing it adversarially resolves some context rot. Whatever they trained these LLMs on will just infinitely double down so I break it with 1 layer of checking and then a judge who looks at facts, since the checker is adversarial.
I know it sounds silly but 1 layer ends uo being way worse than 2.
I could certainly envision a scenario whereby review would increase reliability but not how it would every guarantee 100%, there is a pretty big logical gap there.
In my experience it depends on how much in detail you want to go. Chip manufacturing is a really opaque industry, so in this particular case LLMs might not even have the training data. However, using it for a high-level introduction into something is usually pretty safe from hallucinations.
I don't do animations, but I have an answer. You research a topic well enough to be able to understand if the result is OK or not. Usually it means figuring out some sort of testing.
I'm researching causal inference right now, and my main goal was to make sure I understand how to test estimation on synthetic data.
Basically, it's the same way it works with people. If you delegate a task that you don't understand, and you can't have a credibility proof (i.e. doctors, lawyers), then you research a topic well enough to be able to (1) define the task and (2) verify the end result.
Yep, that's impossible. The hard truth is that most people this lost to LLM psychosis cannot understand that fact. It's better to treat it like someone in a cult, arguing the facts isn't going to help if they refuse to accept them.
I thought after reading the title that the text was about learning something, yet the actual text seems to be about having a system do something for me.
This looks very interesting, something I was also trying to do with my learning.
One thing I wonder is, do you mentally 'fight back' monotonicity of your interactive tool? All seem to be in 3D space, with low-poly, like in a factory moving through the belt and giving you an information + textual description to read more
But sometimes you want to visualize the charts, or graph of simulations, or maybe even the parts of an item in the rocket.
I've had success using the socratic method. I give Claude some topic (say, how the intricacies of the bond market works, the content from which are screenshots of pages from a textbook), and then I go on a walk chatting with it in voice mode. Claude is the expert, I am the student. Claude asks me questions, leading me to an answer logically. I come back with questions, and we back and forth. LLMs arent like they were in '23-'25. I'm almost always skeptical its going to lie, and almost always wrong.
In particular:
- I limit it/encourage it to give me single sentence questions
- I sometimes will ask it to tell me a motivating, human-grounded story, when we're starting a new concept: claude responds "Maya is a bond portfolio manager, and her boss has asked her to quickly price in what happened if yields go down. She knows her bond's average duration, a measure in time, but she doesn't have a percentage, which is what her manager wants. How can she give him a percentage number with just a duration figure and the proposed new yield?"
- I'll often ask claude to let me work through it, to derive the thing myself, often resulting in a string of thoughts with "yes/no" trailers, to get the LLM to reply yes or no only, and avoid derailing my train of thought. If yes, my train of thought keeps going. If no, I've got something wrong.
- I'll sometimes stop and have it craft an artifact. I typically say "build me a Brilliant.org-style interactive demo of the topic", especially when we get into the realm of looking at the actual maths of a thing (for which prose and dialog is not optimal by itself AFAICT)
- I'll do this while I'm traveling, while I'm walking, while I'm doing chores.
It's a stupid word soup on second read, but hey, it works, and that's all that matters with "stochasty the parrot."
"""
In this project, I require a socratically delivered line of conversation. Here's the typical structure to the conversation. I ask some question. You need to factor and reason about how to conduct and deliver a conversation. Best practices would be to limit terminology, or assess with the user whether they have a firm grasp on terminology before you use it. You must be very strict about this, it's unacceptable to just introduce a new concept, actor, phrase or other complication into the conversation without first labeling who what or why it exists for the conversation.
Conversation structure needs to be front-loaded with a brief interview for the user, "you understand X?", "whats your understanding of Y?".
Conversation structure then needs to proceed with single-sentence questions from the agent. User replies with an answer. Sometimes the agent needs to correct the user, but only ever do so with yet another question.
"""
^^ these are the instructions I have installed at the root of a "project".
Keep in mind, this is claude opus 5 low effort we're talking about, in the "projects" area of the mobile app. Here's the process I use to set up its knowledge:
1. I take screenshots of the textbook on my iPhone, and upload a chapter at a time.
2. I have it summarize the chapter into markdown by analyzing screenshots. You could probably achieve this simpler, if you just had the textbook in PDF.
3. I walk with my boy Clau-crates.
I've done this for a couple of weeks and haven't seen it revert back into its typical context-dumping behavior.
On second read, there's probably some clean up I could do. Thanks for making me pull it out and look at it. Things that could probably be improved:
- tell it to cross-check resources online to further ground itself
- use simple, short sentences (long sentences make the brain blur a bit)
- not be sycophantic (it seems like project mode has discarded with my root-level anti-sycophancy prompt)
What’s everyone’s opinion on learning new tech things in this day and age?
My opinion swings between positive and depressing vision of the future.
I still learn new stuff, but I’m afraid it won’t have any value in a year or so.
For example, I’m pretty good at optimizing low level stuff, but right now you can just ask LLMs to do so and they are pretty good at it. They will profile the code and suggest reasonable options like 90% of the time.
They're amazing at it, provided you keep asking the right questions.
Trust me when I say that in the hands of someone who doesn't have your experience, the LLMs would not be getting the results you get.
You might think what you're doing is trivial, it may be sessions that flow roughly, "Instrument this, okay this part is slow, profile this part, OK read the profile output and suggest a better approach".
But your experience will be steering it in the right direction, and you're probably unaware of just how much your experience is doing that guiding, as the LLM shoots off at 100mph, you feel like it's taking you with it, but you will be guiding it a lot more than you realise, and that's where learning and experience comes in, even if you're no longer operating at the lowest depth, your knowledge of that layer will be helping.
If nothing else, the experience to know when something is actually slow is a skill in itself. If a function takes 200ms, sometimes that's as quick as it can realistically go, and sometimes that's literally a million times slower than it could be, and there's actual skill and experience wrapped up in knowing what "slow" looks like.
Except the steering itself is also disappearing, the same prompts from just 6 months ago now need much less steering, AIs are learning to even ask back in certain cases to persuade people with no experience towards the most likely correct choice
Generally speaking, the pattern is that people are overestimating how much "work replacement" will happen, and underestimating how much "work shifting" will happen.
What is fascinating is how you can witness it at so many levels of organization. One example: Employer executive get enamored with moving from labor to capital. They believe that by using LLMs, they can replace a lot of workers. At my place of employment, we have people that are surprised they can't file a Jira ticket describing a product ask, and have it kick off an implementation. You can build the skill to attempt that, but invariably you'll get back questions like "what do you mean by <x>" and "what do you want to do in this case, a, b, or c?"; questions that a product person or an exec are not well suited to answer.
In the past, programmers did that kind of interpretation and judgment call. So then you're in a quandary; who should do that work? Work that previously, you never imagined was an inherent part of what the replaceable code monkeys do at your beck and call?
And then, how do you hire for that? How do you find the training for the people that are experienced enough with... something... to know what a cohesive error response is, or what kind of telemetry strategy is best for that particular product and organization, what collection of product asks are incredibly complicated for what they're asking and can deliver 95% of the benefits at 5% of the work if we just do this instead, and whether you want to aim more towards thick or thin clients?
Who are those people? Wait, those are programmers? Wait, there's this whole collection of inherently human skills that we devalued, by not appreciating they were always quietly doing that for us in the past?
That's just one example. There's a repeating pattern of discovering where the work truly is, work that was embedded in manual patterns we might not have to involve ourselves with anymore, but is yet still essential. So the nature of our jobs changes massively, but the overall level of employment does not.
At least, not in the medium to long term. There is a lot of painful churn we have to suffer through first.
I agree with you in the broad strokes, but replacement by LLMs isn't the only way the level of employment can be reduced. If LLMs can make it so that two programmers can do the work that used to require five, that can result in a very large reduction in employment, even while having (some) programmers is still essential.
Only if there is no further additional demand created by the now lower cost of doing the work. Take lighting as an example. As it has progressed from burning expensive candles to now leds, our demand for lighting has continued to increase.
What sources do you think are good for this data for the recent past? My brief initial search is turning up a lot of contradictory data, probably because the sources are defining "programmers" differently from each other.
I can only speak for myself, so I hope this resonates with you.
I wasn’t even really concerned with optimizing low level code before LLMs and that wasn’t why I was hired either.
However following that low level thread: We can look at the reasonable options and immediately know if they’re reasonable or nonsense. Why? We know the code. Now zoom a level out, where I think our expertise really lies.
Building a complex system isn’t easy. There are customers with requirements, there are budgets, SLAs etc. Sometimes one customer needs X and one needs Y. Our expertise is taking all of this in, and producing something that balances all the different variables. It’s knowing that we’ll expect X events a second so we’ll need Y to ensure we can tolerate failure.
Is it possible LLMs will be able to do all of that too? Maybe. But then why would our customers need the enterprises they pay for?
I sometimes get vague ideas for solving maths/science problems. They never pan out but i can talk in detail on group theory and advanced maths and science topics due to investigating such vague ideas over the years. These days the LLM shoots the ideas down instantly and honestly correctly, i know enough to know "yeah that's right, oh well" and move on. Which actually takes away a huge avenue of learning. I'm pretty torn on the outcome of this honestly.
I'm not 'wasting time' but I'm also not really learning.
I felt this. I had a feeling for a long time that spider solitaire was somehow related to knot theory (legal moves being Reidemeister move equivalents and untangling). Llm shot this idea down very quickly! It was freeing actually.
There's a big gap between being able to ask questions about something and understanding something.
The more things you understand, the higher the chance you'll spot a situation to use them in the future.
I think the best innovations come from times when someone is uniquely able to combine two of their previous experiences together. The more experiences you have in your back pocket the more combinations you have access to and the more likely you'll have a unique combination when the right problem comes along.
I'm choosing to (mostly) switch off the news and continue learning things I find interesting anyway. Maybe the world will punish me for it at some point, but I guess I'll have to deal with that when it happens. The alternative is too depressing otherwise.
There’s still an immense value in training the brain to learn and be able to approach new problems with the sort of procedural thinking that LLMs enable. We can explore topics that we are curious about and develop that sort of “muscle” to continue asking questions when we have them. I have no fear that when those bigger (and existential) problems arise we’ll be well equipped to keep asking questions and figuring out ways to solve them.
So you think building a warp drive is pointless because 99% of work, per your judgement, will be done by AI? Is your contribution meaningless and artifact useless? I wouldn't think so.
It depends strongly on the prompt. Like recently I was also doing some performance optimizations, and if I did not mention profiling none of the LLMs even profiled the code; they merely read the code and assumed, based on their own analysis of big-O time complexity. It is after I explicitly asked for profiling that the LLM started actually profiling.
> I still learn new stuff, but I’m afraid it won’t have any value in a year or so.
This is silly. This would be like arguing that encyclopedias made knowing things pointless. I learn new stuff for me.
Professionally, it's important to know enough to know if you're going in the correct direction. Practically, tokens are going to continue to cost money and knowledge can save you tokens.
I mean, part of the reason the LLM can do that is because you know enough to direct the LLM to do so and verify the results to some degree, right? It's good to learn new things because:
1. It satisfies you curiosity (and curiosity is always valuable)
2. You can better utilize the LLM to expedite something you now have knowledge about
3. You still improve as an engineer/programmer/prompter/whatever
I still think it's very important not to outsource everything to AI because there is a lot of value in learning and doing things yourself which is an important part of life.
I have staff ranging from 10 years of IR experience to right out of college.
I can tell you that there is an enormous gap in ability between them despite them both using LLMs for daily IR work.
The reasons aren’t complicated. The senior responders have tacit knowledge of how breaches evolve and what to look for which gives them a much better framework for where to employ the LLM.
The juniors will normally start from “here are some logs, look for weird” which is fine but leads to tunnel vision and a lack of confidence in their reporting.
I don’t mandate that anyone do work with or without an LLM. I hire seniors based on experience and juniors based on interest. But my experience has so far been that our best up and comers focusing more on learning the technologies instead of leaving those details to the LLM are developing their intuition and understanding faster and in a more robust manner.
The biggest thing I've learned from doing stuff like this is that there are no shortcuts. At some point or another, to truly learn something deeply, you've got to dig in to the boring details and do things the hard way. LLMs can help with this...but I find it's usually tempting to try and just offload the boring stuff to them, which doesn't work.
Having an initial higher level understanding across the domain is extremely useful to contextualize the deeper stuff. I think boring details is a very leaky characterization, but I'll continue with it.
In my experience it's infinitely easier and faster to learn deep, "boring" things when you understand how they relate to your shallow and wide understanding of all of the related components.
The LLM is merely a tool. And you can use it for domain discovery that enables efficient deep learning at an unprecedented rate or you can develop a cursory understanding of a topic and think yourself an expert.
I’ve found that if I’m not struggling I’m not learning. If stuff is coming fast and easy that’s a sign that what I’m doing is not stretching existing skills enough.
It’s true of most things. Running, dieting, weightlifting being uncomfortable is a sign of progress.
> The biggest thing I've learned from doing stuff like this is that there are no shortcuts. At some point or another, to truly learn something deeply, you've got to dig in to the boring details and do things the hard way.
This is the only path to mastery, or understanding if one prefers. There are no shortcuts to a person achieving deep understanding (a.k.a. "Aha!" moments).
Can a tool such as GenAI be beneficial to someone who already has done the work to understand? Absolutely. But it cannot infuse mastery into a person simply by its use.
Only the time and effort a person devotes can do that.
Quite - you still have to do the work yourself. I think LLMs are best placed to act as an eager tutor that doesn't mind discussing a topic ad nauseam until you're certain you understand it.
yes i've found that there are a few topics that i've really been able probably 10x my understanding of using LLMs, in particular in getting me over hoops that are hard to navigate when solo, BUT I have to be really careful for it to not just show me the answer all the time.
I’ve been using LLMs to create readable rewrites of RFCs and specs that interest me. It is not precise enough for implementation use, but it has increased my understanding of the underlying RFC.
Another useful approach has been asking Codex to implement complex things, like a Kademlia DHT or BitTorrent client in a literate style with the explicit purpose to increase understanding by reviewing the source code.
I usually learn compelx topics using the image generator of an LLM, capabilities like GPT-Image2 or Nano-Banana.
I give it like a complex paper -> turn into visualzation or a poster, then ask quesitons and it helps me understand someitmes a very complex paper rather easily. The jump that nanobanana / gpt-image-2 had done is pretty wild.
The title is not representing what the post is about. “Use LLM to learn complex topics” here actually means that the author asks an agent to describe the problem area, and then implement a simple web-based simulation game, and by playing that game, the author actually learns about the topic and its constraints. They use chip making as an example.
Its fun but is it really effective ? I mean I checked the LLM one and I came out more confused about a topic I already know about, I find the best way to to learn using LLMs is to just generate an example try to somewhat get a mental model of how it works and then ground my understanding with traditional documentation and resources, its an iteration of a technique I used to do in college where I would read the textbook questions first to understand what is important and then read the chapter
I assume different ways of learning work for different people. For me personally, it's taking a piece of paper and drawing the diagram of how things work together; of if it's some math, then, again, using the pen and paper to follow the text. I can very much accept that for some people playing the simulation is a good way to touch the new problem space. I can easily imagine that for some topics, let's say, traffic signal automation, a careful simulation game will probably give more information than reading papers or manuals.
Game-based learning, described by Comenius, works if someone else prepares “a game” for you. E.g. like a dungeon master. :)
Otherwise you probably get more confused as you have mentioned.
On the other side, Peter Diamandis describes a situation where a bunch of kids were given a internet-connected computer and they had no teacher. Instead of it there was a “grandma” that checked kids from time to time.
After that there was a knowledge test that revealed “no teacher” approach was more efficient.
Thanks! The original title was "How I use LLMs to learn...", but somehow HN removed the "How" part. I even removed the initial post thinking it was a typo on my end and tried to post again, but I stumbled upon the same behavior.
I only use LLMs for explaining things if the subject is one I’m deeply familiar with and can independently and easily verify the facts. Doing so with new knowledge areas is very risky for obvious reasons.
For example, “explain how the code in this file works,” I am familiar with the overall codebase, I know the purpose of the file, and I can read it or write tests to verify if I suspect what it’s telling me isn’t correct. Or, if it’s really important, I can overcome my introvertedness and ask the team member who wrote it…but that’s a last resort nowadays, which I am very thankful for. In 99% of cases since at least Claude 4.2 days, Claude and Codex have been very accurate. Gemini on the other hand messes up more frequently and sometimes does weird things like try to delete files it’s not familiar with, at least the 3.6 flash model I’ve been using lately does this. But, code explanations are still good for the most part.
"Complex topics" in this case means reading 22 AI-generated paragraphs that supposedly cover the entire chip manufacturing process. If the author seriously thinks this level of detail is complex then they have psychosis.
My main issue with using AI as a learning platform is that unlike documentation, books, Youtube videos, there is not really a process of having someone "review the learning material". For example, I can always read the review of some book or ciriculum, the comments under a video, or if it's some for of open source documentation you can check the PRs and verify to some extent it's claim. With AI I can't really say what it has halucinated, because I am learning a new thing, I don't have that benefit of previously reviewed material.
This is (exactly) why I very strongly tell people not to teach themselves with an LLM. Particularly from the ground up. If you do not understand the domain, you cannot learn from the model because you won't know what questions to ask and it certainly isn't going to answer all of them for you.
Of course, and that is what I do, but it becomes and additional mental and time consuming effort. As an dumb example, if I am learning a new programming language and trying to grasp some concept, I'll check the docs, find the section and read it. I trust that the source in the docs has been already vetted by other devs and the authors. But if I ask the AI the same thing, I then need to verify it's claims usually by asking it to check if the info is true, and then possibly opening the source link (lucky for me Claude provides the links in the desktop app as footnotes). It's not a question about AI, it's about the trust I have of this tool, it builds over time, but as soon as the AI makes an assumption or a hallucination we are back down to square one.
I assume this will become less of an issue in the future as there is more trust between the AI tools and me.
Wrote about this awhile ago, and it hasn’t changed:
> Time and time again, when talking to people who rely on ChatGPT, Claude, Perplexity, and other general AI tools, I hear them say, “AI is incredible. It handles nearly everything I throw at them.”
> “What does it fumble with?” I’ll ask.
> “Well, it still gets things wrong when it comes to my line of work.”
Buildings are a fantastically complicated and interesting complex problem space. I found my architecture studio students using it to ask technical and code questions about the buildings they were tasked with designing. I observed that the inaccurate information it was returning was compounding...
To be clear, I say "inaccurate" rather than "wrong" in this case because even if the information it returns is factually correct to the question being asked, students don't have an understanding of the complexity of the interdependent tectonic, regulatory, and spatial / experiential factors of a building sophisticated enough to ask their questions of the specificity and nuance necessary to get a good output that addresses the entire problem.
Anyway - with the students still learning to ask questions the right way, and the conditionally-incorrect facts making their learning more complicated rather than less, I hit on a strategy for them to use LLM's that seemed to help much better.
I suggested that instead of ask the LLM for the factual answer, or even better for the facts and an explanation, that they ask it to direct them to the proper place in the source material to find the answer themselves. Then, to treat it like a lab partner. IE:
Hey Claude I'm looking for "x."
Claude: "look at foo, bar."
Thank you - chapter (foo) part (bar) table (goo) says "car." However I notice that footnote (hoo) says there's an exception if "dar." Which is what I have. Walk me through this exception...
It seemed to have good results as a guide to understanding the disparate bodies of knowledge that they will eventually have to keep together in their heads and work synthetically and non-linearly through, rather than just as an external source of blindly trusted authority.
The author says it's "100% accurate and free of hallucinations" but I am sceptical. Although I don't know much about chip production, when I ask LLMs about advanced topics like memory order LLMs tend to hallucinate more than entry-level questions. But the point is that I didn't found that LLM hallucinated before knowing it deeper. It's possible that LLM hallucinates but you don't find out because you are just learning it.
These animations are great. I have learnt so many things from LLM's, cross checking things is easy enough, but the hallucinations are really not much of a thing any more (in my experience). When you deep dive on things it does seem to get very wordy sometimes (this is claude anyway). Its taught me flutter and dart without opening a book (with the occasional reference page), set me straight on monads finally, refined some linear algebra, various bits of history, and philosophy, I'm learning spinors at the moment. Is some of it wrong - maybe, but it's not like my brain is 100% accurate any way, and when I need accuracy I look up references. It is fantastic getting a broad overview of a subject you don't know or a precis of a current subject, and its so much faster.
Hallucinations bring to question what you think you've learned. That's going to cost long-term if you labor under mis-apprehensions until you maybe figure out you learned something wrong.
where's it mentioned on that link? I can't seem to find hallucinations (don't worry found it).
Yes, sometimes its wrong, most times its right, cross checking is fairly easy, not using it because of the possibility its wrong seems a baby with the bath water thing.
I learnt a lot of functional programming from it, stuff I've always wanted to learn, but just didn't have the time and really the sources can be difficult, it really explained things well, and as someone else said in this thread, you can ask questions over and over until you understand, asking a person that (if you can get an expert) would drive them nuts. Maybe my experience isn't typical, its hard to tell, everyone reports something different.
having llm audit all my work/code, and generate review pages (as a teacher) of before/after with working examples is incredibly useful. Its something no course can do for me, even a teacher wouldn't have enough patience to go through each one of my mistakes.
having things defined/have correct solution to compare for review is useful, and keeps llm on track. Don't think i would trust llm if it were reviewing it all on its own
Imo the podcast and the video are better served as background material for some other task. The low bandwidth becomes an advantage because it's often ok if you miss out on some parts due to lack of attention.
Indeed it's often a waste of time to just focus on talking people fully if you want to learn fast, reading and especially deliberate practice are better for that. But if you don't have the time, energy or focus, then listening to interviews in the background can be useful supplementally
I am very interested in figuring out how people use LLMs for learning. I definitely have the knowledge, but it is severely autistic in a way.
OTOH, I am curious if there's a "practical value" to this exercise? If the LLM already contains the information and implementation knowledge to implement the networking stack inside an FPGA by itself, what value do I gain by learning about HDL, TCP, the bespoke Xillinx tooling, reading the documentation, reading papers on the implementation and going through every bit of details and theory.
I feel like there's a meta skill that is more worthwhile for "practical value".
I’ve written a skill that I basically feed what I’m looking to do, some ideas I had for accomplishing it and any other details like tech stack, etc.
The skill then riffs with me, judging my ideas and suggesting alternatives. We go back and forth until something useful comes out of it. This process isn’t unlike how I do normal development.
However, once agreed it breaks the work into “steps”. It then creates a tutorial for me, for those steps, explaining each line, why each change happens etc. I can then ask questions, muse about an alternative idea etc. Then I do the steps, and I’ve learned and gotten what I wanted to get done.
This has been how I’ve been learning Godot and making a game for the past month or so. I didn’t go in blind, I started with a course from GDQuest so I could feel confident guiding the tutorials. I will say though, having a tutor to bounce ideas off of has been really useful.
I still try to figure it out myself, consult the docs, discord etc. But if I’m stumped I’ll run my tutor skill and have some fun.
I’ve started doing a similar thing after reading a post on hn about manually applying the code so that you actually understand it.
I have done this for all my work this week and it works quite well.
For one it lets you actually query the LLM as to why, their plans give a high level not every single change and it allows you to correct it as you go and the plan will change.
I've been using them by reading some docs/wiki/tutorial, then when I think I understand something trying to do a rough explanation to the LLM and ask if I'm right. I'm usually making some analogy to something I already understand a little. I'm usually partially right but missing some key bits at the first pass. I go back and forward asking for explanations of various bits or asking for resources around the area I'm not understanding. Often times just discovering the relevant name for the area of study opens lots of doors. I basically use it like I would talk to a knowledgeable and patient teacher.
As for how useful it is to understand thins, I believe it's still useful and hope it will continue to be.
For everyone who thinks you can’t learn with LLMs, Dr. Cat Hicks, psychological scientist and author of the recent book “The Psychology of Software Teams[1],” who worked at Google and founded the Developer Success Lab at Pluralsight, has written two skills called learning-opportunities[2] and learning-goal[3] that use validated learning science to help you learn while using LLMs.
Cat also has an awesome podcast with her wife, Ashley Juavinett, Phd, called Change, Technically.[4]
I encourage everyone to check out her work! She’s dedicated her life to helping software developers get the support they need inside organizations to be seen as humans, not just robots.
I use it by telling it my background, giving it a rough timeline and asking it to create a learning timeline, save progress along the way, and git push / pull periodically so I can use the same thing on both Linux and Mac. I tell it for each phase in the learning timeline, present me information, then challenge me on it. If it's code, it challenges me with a coding challenge, where I use it in an IDE plugin. If I'm learning something that isn't strictly code, then I ask it to give me info, then challenge me with questions and grill til I get it right. I ask it to save what it think I struggled with, so that later we can drill it again and I can also review it in an .md file.
Does anyone else use Claude like this?
It's sped up my learning by 10x. I struggled with 'just reading a book.' Take kubernetes. I hemmed and hawed and spent years periodically reading some dry book or blog or official doc, falling asleep, and forgetting while I got busy. Now I'm aggressively working with it, almost like I'm addicted to a gamification, of getting through our learning timeline, and I'm excited to move forward as quickly as possible and pass its tests.
It's like a fake teacher, because I can also ask it to drill into a topic or re-explain itself if it made no sense.
The only thing that worries me is, sometimes I'll say something like, "Um, are you sure about that?", and it'll apologize and correct itself. I barely challenged it!
The little tool it outputted is nice, but click around the stages and the text is not high quality at all. The snippy titles, abbrievated explanations, I wish a few more iterations and thought was put into the actual main textual content. Especially for 'complex' stuff
If you're using LLMs to learn or for research, and at some point you don't end up engaging with an actual resource (books, papers, lectures, web pages, etc) then you're playing yourself.
Sorta. If its response is grounded in actual material, and you're thorough, it's not so risky. As with all learning, trusting one source is a risk in itself. Hell, I didn't even trust my physics textbooks in college. Physics.
I use LLMs to learn deep technical concepts. I really like them because I can spend countless hours a day understanding things and building an investigation file with all my findings. I code examples and test the findings. It has helped me understand basically anything.
I'm using LLMs right now to build a terminal browser, a GUI browser, and a PyTorch/LibTorch replacement. It's really fun to be able to learn and make progress this way. It's like reading multiple interactive books, where every concept can be explained again and again until I understand it.
The main bottleneck as an engineer is no longer writing or testing code. It is how long it takes to understand complex systems. This is a really nice approach that, if you have the tokens and the patience, feels like I great way to learn something and I think we'll see more and more stuff like this.
I had a similar realization a few months back and am working on a tool that generates "mermaid walkthroughs". It is 1000% less pretty but it is fast and is pretty good at explaining how services work or what a code review does or just as a way for your agent to explain some decision to you.
But AI will be better than you at those topics as well, and when someone needs an expert in that topic take a guess who will they approach in such scenario.
I don't think we are even that far when the complexity AI can handle surpasses 99.999% of what humans can handle, where AI make e.g. physics discoveries beyond the grasp of most humans and it will have to "dumb it down" when talking with humans -even physicists- but not with other AIs
I just gave a simple prompt to Kimi k3 when it came out, and then just using it and asking it to improve based on my own taste/gaps etc so it’s a long drawn out process. I’m doing that for the first one, will iteratively keep improving it myself as I run through it. Maybe will post on HN after that.
Gamifying the presentation could make topics more accessible to others. For me the overhead wouldn't help with my own learning. Also I've been burned by just learning things mechanistically (e.g., coding, applying algebraic rules), so I'm leery of learning just by making flashcards or models of the topic.
I find LLM's do great for learning when I ask what are the principles, how the main applications work, what are the key drawbacks, where are the growth plates in the field, etc. - the kind of thing a good advisor points to. Sometimes I have to ask it explicitly to use topological order of topics and show relations, which often highlights the gradient changes in the learning curve. For pruning, it's surprisingly good applying philosophical heuristics - Occam's razor, or Derrida's differance (the difference that makes a difference), etc.
And finally, no learning is effective without problem sets, and for those LLM's at times get me over blocking issues.
The degenerate case is memorizing the glib phrases regurgitated back to me; they're helpful and functional enough to get me into real trouble!
> In plan mode (using CC, or OpenCode) I ask a model to build the foundational knowledge for X topic. I ask it to review the accuracy of the knowledge base it built in the previous step.
> What you get is a beautiful animation that is 100% accurate and free of hallucinations.
I wish there was a LLM tool to explore a topic recursively, as a tree or a mindmap. You would start with some high level concept (say "cryptography") then dig further and further to more specific topics.
I think that would be a way more natural way to explore than being stuck on the classic linear output of a LLM.
Something that I have realized recently is that it has become so easy to get an answer to almost any question with the help of chatbots that its almost unnecessary to spend any effort thinking about the problem or the solution. I feel like before when I had to spend time researching a problem to find an answer I learned so many things around the topic itself which helped me understand the problem itself better and gained a deeper understanding. Today it feels like you can have an answer to the most complex questions you might have, yet you gain a superficial understanding of the topic and might forget about it quickly.
Personally I'm excited about these sorts of experiments. We all learn in different ways, and these sorts of techniques allow us to create "on-demand" syllabuses and lessons that fit our learning style and learning level.
It's not perfect, but I'm optimistic this will be a useful way to teach/learn in the future.
And, to be clear, I think this will be best utilized within a group/community setting. I don't think it will replace teachers or classrooms.
This guy is severely milking it now. If learning means building an inaccurate and incomplete understanding of the topic then go hog wild. Otherwise https://news.ycombinator.com/item?id=49209049 sums up my feelings about the author's attitude.
haha, I did the same to study new topic and then realized how personalized education will turn out in the near future. The only thing left is keeping a curious mind and ask a lot of questions !!
Really neat idea, I think it is one of the best ways to exploit the combined building and explaining capabilities of LLMs.
I am currently building an app/game to explain friends and family concepts around wealth management and wealth building. Games are a great way to hide complexity while still including it in the « guide » you are making.
i looked at the animations, they look cool, and i don't think i will enjoy learning things that way. as someone else said, there's a lot of content already produced on these topics. i also think the level at which these animations are playing, they are actually hiding the 'complexity' of these topics.
I tend to ask the LLM for a single HTML page explanation, with a pedagogical approach. Something about dropping the word pedagogical leads to a more structured outcome, but I haven't quite figured it out why yet.
Very cool, I like the visual learning nature of this and the auto play once starting. The game graphics are engaging which counts for a lot these days, I feel my attention span suffering after using agents for the past year.
I've been working on a similar process of pushing to github pages, but focused more on having "practice sessions" with coding blocks to test content. Using webassembly and mock servers to mock backend endpoints Here's one I built to build a full stack llm chat system in the browser.
> What you get is a beautiful animation that is 100% accurate and free of hallucinations.
How do you know if you're learning this for the first time? Very risky to learn from LLMs. I've done it, but you have to keep your wits about you. Lots of "oh of course you're right - what I just told you was completely wrong".
My favorite way to learn infra topics at work right now is asking for a humorous analogy involving monkeys and bananas. I tend to remember the result, and it gives me reference points for new topics.
LLMs can help you understand a language, but they can't replace learning the vocabulary. Words and phrases still need to be learned the old-fashioned way: repetition
I don’t think I really agree with the author’s approach here, but I will say LLMs have been a huge help to me as I’ve been reviewing linear algebra and diving into signal processing. Anything in a textbook that I don’t fully grasp or am confused about, I just take a snapshot or copy paste then ask a model to derive it or explain it in different terms.
It reduces friction a ton, but at the end of the day I’m not skipping anything.
The selected topic (chip manufacturing) is being presented at a shallow level at best. Go read Wikipedia’s article [0] and see how deep and broad the topic is. It also skips quite a few steps, and utterly glosses over how insane of an accomplishment EUV is.
This is my biggest societal issue with LLMs: they allow you to think that you’ve “learned” a topic because you read a lot of technical terms. I don’t think it bodes well for the future.
YouTube has so many truly wonderful videos on chip production. I admire your approach but it seems like a lot of people are in this ai maxxing phase where they reach for ai for everything despite their being ready, high quality things already available for free
I'm also not sure you can really learn chip production from widely available public information. It's a hugely complex industry where the details tend to shape larger strategies. For example, you can't really understand the relationship between Micron and TSMC without some awareness of the trade-offs of memory processes for peripheral transistors.
YouTube is just a big dump of information. Having a structured way to learning along with interacting helps you learn. Otherwise you're just binging information
i've asked llms to write a presentation for me on a given topic. for whatever reason, that finds the hidden layer responsible for explaining things well.
That seems like a terrible way to learn. It’s a neat animation but cmon, there’s like educational TV programs from the 80s that explain this so well, in Germany there’s “Sendung mit der Maus”, not sure if they have a segment on chip manufacturing. But in these clips you can at least see the real stuff instead of some half wrong animation, just let the LLM write a few paragraphs for you or better find an ACM article or book on the subject, probably still takes less time than coming up with that animation…
God does everything have to be productized and glorified as if you’ve invented a new way of learning. Read some books!
I had an internal company assessment I needed to pass before end of our fiscal year. The study material consisted of 10 ppt decks about 80 slides each (so around 800 total). I had an AI read all the decks and compose a study guide with quizzes along the way. It came up with a 100page word doc that I used in place of the decks to prepare. It worked very well for this including, like you said, quizzing me over various sections.
(Yes I confirmed it was ok to use AI with the material)
This animation is worse than useless. I've taught a lot of students. This is not learning, it's stamp collecting.
You're just memorizing a nonsensical recipe. What are the constraints? Why do we do X rather than Y? How does a particular thing scale? etc.
All you're doing is fooling yourself into thinking that you've acquired some knowledge. When in reality you haven't even learned the basic mental model to reason about this stuff.
You've learned something when you have a mental model that makes correct predictions. Until then you've memorized it at best, and as with most memorized things it will decay exponentially and will be gone from your memory soon enough.
i had Bolt create an interactive app teaching me Oberon language but despite several attempts to fix issues, there were still glitches in unexpected places.
So the experience was... meh.
Then I went back to a book written by humans (Eric Nikitin's "Realm of Oberon").
> I personally find the style used by LLMs to explain things difficult to follow. It's just too simplistic
I also struggle with LLMs explaining things, but for the opposite reason.
I consistently have problems to get short, precise but plain/simple answers.
Instead I'm overwhelmed with walls of texts, often filled with jargon that is a mixture of imprecise and unneeded.
The style at which I learn better is by asking about stuff interactively. I ask you what something is, you give me a 3-4 sentences top answer. Then I explore and dig into the topic from your answer on the things I want to know better.
I am greatly disturbed by the idea that any of the demonstrations involve "complex topics". Whether it was the silicon workflow, LLMs, or EUV, I saw nothing complex. Are these supposed to be freshman undergrad or even high school level explainers? I suppose I didn't see any glaring errors on a quick glance, but definitely lots of details were glossed over.
Perhaps basic special relativity could be done this way, or simple derivatives, but definitely not general relativity or integrals, much less PDEs. I guess I was hoping for a 3B1B type output. Oh well.
And then I begin to think to myself that I should just read a book on the topic written by a trusted source who put a lot of effort into teaching the topic properly and presenting the information in a thoughtful way. So, I am back to books and mostly try to use LLMs to clarify certain questions or ideas I have.
human conceptual thinking is very much a multi-dimensional graph, which relies on light "approximate" concepts that are "good enough". LLM AR token generation is extremely one dimensional and doesnt care about the "weight" of the concept behind a token.
LLMs hold billions of parameters in "mind" at once. humans hold like four "concepts".
This is the essential mismatch and the primary reason LLM conversation can be so painful and exhausting.
Explaining this and limiting "concepts" to four at a time tops is one of the very few AGENTS.md / system prompts I always use, and it has proven invaluable time and again.
Thinking traces show how effective this is at forcing the LLM to simplify its thinking.
[edit] Also, myself and nearly all of my peers are struggling to choke down the flaws of LLM tooling along with the benefits. the speed at which LLM adoption is being forced, without truly crafting them into quality tools first, is not ok, and not normal.
LLMs have stirred an inhumane hunger and fear. the tech is fine, but the way tech companies (creators and consumers) are behaving should be deeply questioned.
it's NOT normal. it's not ok.
The full PDF is worth a read (Figure 1 may be of interest to many here): https://www.cambridge.org/core/services/aop-cambridge-core/c...
If "attention is all you need" then it's something we do indeed lack, in comparison to LLMs! But it's an interesting question: might machine cognition benefit from similar bottlenecks in an attention algorithm? Advancements like Kimi Linear seem to indicate that we're far from the finish line: https://arxiv.org/abs/2510.26692
For example having an LLM summarize a dense topic and to find books so that you can filter faster and spend time reading those books works way better than having the LLM summarize the books or the topic (or even relying on second hand information). Another one is having the LLM quiz you on your topics of interest. With questions tailored to attack specific areas that you struggle with. Its wonderful at this, nothing I've used comes close to what an LLM can do here.
You define for yourself what your goals are, slowly refining them as you learn more, and use LLM as a tool. This ,I find works best for learning.
It's long been the case that the best way to learn something is to teach something.
Which is pretty unfortunate for those that want to learn. I used to enjoy writing documentation at work, it was my favorite part of the job. And it did feel like it benefited me more than it benefited all the people that were (or weren't) reading my documentation. Now I can't really justify spending much time on docmentation when LLM's can do it in a fraction of the time and it's "good enough"
And you notice when it's a topic you know well or something like software where you can immediately tell the options it's giving you don't exist on the page. Leading to the amusing statement "LLMs are bad at what I do but great at everything else".
The question: what's the net positive gain of turning people who know nothing in a given field into sub-novices, while weighing actual experts down with work slop and marginal returns?
And I wonder what the true cost is of arming so many novices with that level of dangerous knowledge.
Tangentially, but related: I'm old enough to remember when the spirit of your comment was pervasive on HN.
I’m old enough that I worked my first IT summer job the same year slashdot was founded. I’ve seen a lot of tech tribalism form and dissipate, and this one didn’t feel organic. My gut says a lot of the us-vs-them tension originated in a deliberate campaign to cast AI boosters as the tech industry in-crowd, and ‘other’ the people not on-board. Who knows.
Just hoping folks don’t get hurt due to people not understanding what they’re doing with these things but believing they’re competent.
The body of literature on learning theory, and beyond that on specific types of learning and specific mediums such as learning from text is so rich there are way more useful models to draw from. Believe it or not, prellm, researchers in the textual learning field had already demonstrated you can achieve performance equal or better than novice tutors using pretty basic computer aids that follow specific hint/pump interaction structures. Guiding an LLM to use these findings has evidence backing it and is way better than telling it "i guess be like socrates". The problem is, to realize there might be richer more effective and highly researched ways of tackling the problem beyond the first fart of a thought you had one afternoon requires the deep respect for expertise and specialization that precisely basically everyone in the AI space right now fundamentally lacks.
https://adaptive.bounded.cc
Trying to diagrams/animations didn't yield good results even with frontier models. But pure text, any model does a decent job.
So it may be very slow or become unavailable, back end can't handle that, no caching whatsoever.
I'd imagine an application that uses LLMs will be created that better manages learning. It's just not clear what that UX is yet- it's obviously not just a chatbot
i run into context window limits, or practical limitations of digitizing the book
I even wonder if this behavior is due to next-token prediction architectures, somehow.
I know you probably don't consider it dense but wondering if someone can shed insight.
I find them like empty calories, like programming youtube tutorials. They maximize for feeling learnt instead of steady progress
Personally, I find that its generated prose tends to have an undue weight to it, almost as if every topic I ask about somehow bears a heavy burden, or is otherwise load-bearing, to use its parlance.
Quite puzzling, really.
I have a personal theory: LLMs are *fundamentally* handicapped at perceiving what's going on in the mind of the human (this can't be "innovated away") and that's at the root of what makes them suck at conversation.
Next time you're chatting with someone, notice how much understanding is shared without anything being said. E.g. the other person might share something deeply disappointing, and they can tell without you even saying anything whether you get what they're going through. This unspoken-yet-communicated information guides the conversation. Or as another example: humans can read the room -- you walk into a room and immediately adjust your demeanor based on what you see and sense.
LLMs are totally blind to things like this, and this adds an inescapable awkwardness to interacting with them. I don't believe they'll ever grow out of this. Which thankfully implies more long term demand for humans instead of robots. :)
Does anyone have a read on if this is primarily a Claude issue, or if all LLMs do this?
I think this is one reason why LLM text is pretty exhausting to read for long stretches.
It's possible that this quality you describe stems from the extensive training corpora utilized by the major AI labs. These almost certainly include work from the esteemed economist Jacob Silj:
https://www.youtube.com/watch?v=Poc1upTejD8
I view LLMs in education similarly to office hours. Some people abuse it to get homework answers without grappling with the material, but the optimal amount is not zero.
LLM certainly not a replacement for a book, where you get someone’s extended personal approach to a topic, thoughtfully organized, reviewed and edited, often times actual courses taught based on it, with answers checked and errata available online.
Perhaps the best example has been a native macOS app that is a completely custom text editor with built-in debugger, lsp support, fuzzy finder, etc stuff you'd expect. Inside the same app is a library of books i can read within the app completely formatted and for every chapter/section of each book that is a quiz to take (LLM generated of course), a "recitation" tab where i am asked a question and say outloud my response to the AI to evaluate me on and then finally practice problems to do within the custom text editor (these are usually programming books). The reader also has ai re-write built in.
As neat as this is, and i worked through K&R like this, i have ultimately fallen back on "just read the damn book and go to the AI when you've got questions."
> I get exhausted reading LLM prose
So much this! If I see one more sentence with the words "genuinely" juxtaposed with "load bearing" my head is going to explode!
btw, I am building the tutorial here for anybody interested in this topic: https://github.com/avilay/learn-probml
it also researched vision correcting displays for me and i can finally put that idea to bed - i was never really going to pick up an optometry textbook tbh. plus it was able to pull together a bunch of geometric and physical context about light and the eye plugging exactly my personal knowledge gaps.
in general i suspect these materials might not be that interesting to others because they are so custom to my learning style and personal needs and preferences.
these are usually not one shot documents but rather many prompts deep before i get something I’m willing to sit down and read or study. but dramatically quicker than assembling it myself from primary sources. i wouldn’t say it matches master expositors but then they’re not available to write on any topic i happen to need right now.
plus I’ll just have a live voice discussion with the system when i go for a walk and there are still things bothering me on a topic. it takes a little patience but if i’m in the mood it’s amazing.
i generally find that it can help track down specific references if i suspect hallucinations. but especially on factual topics my experience so far has been extremely encouraging.
But even with Claude, it's it's really the prose getting in the way you can install the caveman plugin or tell it to use that "standard technical English" thing.
Any more detail you can share? Do the others feel more "human"? Are there any that are particularly digestible/human-friendly?
I've been wondering for a while if this is just Claude because I mostly use Claude, so this is very telling.
I tried using a new agent service recently and could tell immediately that it's powered by Claude due to the way it writes.
So I have the LLM offer a very short explanation of something, and from there's it's just me asking questions. Anything that feels fuzzy or not fully internalized is something I poke at until I'm satisfied.
It really has helped me develop a sensitivity to what I understand vs what I don't, and the ability to drill into any part of it is amazing.
And yes, it is not that it is just presenting the facts. By me taking control of the direction the questions and answers go, I can flesh out my mental model. I won't retain every little thing it tells me. But I am much farther ahead than before.
I will say, opus 5 is an egregiously bad case of this, but other LLMs have this too, just less bad.
Agreed.
I find Opus 5, and even Fable, to be overly wordy in eg PR descriptions and code comments.
However, I suspect that's more to do with what they are trained to do by default than LLMs in general. I have a little setup where I tell Claude to work together with Codex to tighten up prose and comments, and for me that produces much more palatable text that needs less human editing afterwards.
That's speculative, isn't it
While my advice is specific to learning about codebases, the way I do it is to have it generate mock data and put it in the local development environment, and give me some exploratory commands, and then ask away. It's a machine after all, so I don't have to read its preceding prose to understand whether it did tell me something, it can just repeat it however many times I ask it, and the hands on commands etc. give me something to actually try and implement.
I'm not sure I'm better off with humans though -- I'm not qualified to judge whether a source is a proper authority, not an I qualified to judge whether someone knows enough to point me to a reliable source.
It seems this is a fundamental epistemological problem to which there may never be an answer.
This is my biggest gripe with reading AI-generated text as well (ignoring the meta issue of whether it's worth taking the time to read something that an author didn't think was worth the time to write). It's gotten to the point that weird AI-style analogies just take me completely out of the text and kill my interest.
And I can usually tolerate a lot of purple prose.
It's a loop that uses adversarial review to check several dimensions of the writing:
https://github.com/Vibecodelicious/llm-conductor/blob/main/w...
I’ve found the tone of Kimi K3 to be less obnoxious. Unfortunately it doesn’t wholly solve the issue, I don’t think any LLMs out there have a truly pleasant writing style, but at least not every assumption is “load bearing”.
Like I said, I'm essentially continually prompting to refine the material. LLMs certainly continue to append, and never cut back. It just keeps spitting out additional content at me. So that's a bit annoying too. But I can basically get figure out what's going on with a few extra promps.
If youre curious what i've got so far... just be warned it is quite literally AI slop plus me continually prompting for clarification/cleanup etc. : https://github.com/cmoscardi/ai-for-ai
It generates tutorials for you, and serves a webpage that lets you complete them. It does a remarkable job.
It still has a bit of the LLM prose problem, but it does help you fine tune the ‘voice’ it uses.
I've stopped using CC because of it. I find it insufferable.
This problem doesn't get talked about enough and is second only to the hallucination problem IMO.
AI produces so much noise to wade through in order to find signal, and the more expertise you have in a field the more that costs. That noise directly subtracts signifcantly from productivity gains.
And, I think the problem is directly related to the hallucination problem. It feels very much like an effort to kitchen sink the response in order to provide some value among possible hallucinations.
It also seems to be a byproduct of Gen AI operation. It just fundamentally doesn't understand what it's outputting, so doesn't know how to narrow down to the most salient bits.
It's just long. It just doesn't shut up. It's overly verbose. And you can't tell it to be concise or you degrade its quality.
If I ask what an integral is, the correct answer is that it is the continuos analog of a sum, generally used to calculate areas and volumes.
It should really be a single sentence, and then let me ask more about the terms I don't understand, and here's the beauty, in the previous one there can be only 5 terms I cannot know.
An LLM will vomit an entire page or more of explanation which isn't bad per se, but is an answer to something different: "give me a short introductory explanation to integrals". And that's not what I asked.
Try it out, fairly sure that if you out in 100 random words for 30 of them it will just refuse to translate them (it will copy paste the original word into the target language) or it will do silly things like use the target 4th dictionary definition instead of the primary one).
I'm actually going to make a prediction here as well. I think you will soon realize that using LLMs to clarify certain questions or ideas you have will turn out to have frustrations as well. And that you will soon direct those questions to either peers you know in real life or internet forums which are very likely to have a non-AI policy.
Much more likely that people will believe themselves to be an expert in a subject after having had a conversation with Claude about it.
Often "be concise, to the point." is enough, but you can also paste it some stuff you like as an example text and ask to do style transfer.
Literally saying "one sentence response" solves most of this problem.
I'm not sure I follow how this is actually guaranteed? The fact-checking process mentioned just seems to involve asking AI to review its own work.
And even if they are related - if Opus 4.8 always has a 1:100 chance of a specific hallucination - then running the same model twice does indeed dramatically reduce the odds of an error in the final output.
Personally I think this is a bad characterization of using LLMs to fix up LLMs because while you can never guarantee results this way (as the quoted line claims here, which is worthy of criticism), it is, in practice, useful to use LLMs on top of LLMs. And there's no infinite regress. Auto-mode in Claude Code, for example, seems to me like it's been successful at making the system more safe than --dangerously-bypass-permissions without prompting the user for permissions constantly.
What triggered my response was the “just review the output with another LLM and it’s perfectly correct”
Full story in the book
People don’t even have to be lying to be wrong about this stuff. Someone can learn enough about a topic to be halfway up Mt. Stupid in no time flat, and in doing so, think they not only truly understand the topic at hand, but might be particularly adept because they were such quick studies. People that know less are impressed, because why wouldn’t they be? Anybody that knows more than them sounds like an expert. And people that know what they’re talking about cringe at the overconfidence, and probably try not to engage: who wants to have to prove that someone’s boundless confidence is entirely baseless? Most of the time, they think the actual expert is full of shit because they think they’re the expert. It’s incredible how many times I’ve had people in tech confidently, even smugly “explain” design concepts and strategies to me that they did not actually understand, knowing I was an experienced, degree-holding designer… and they didn’t even have a chatbot’s lips on their ass telling them how smart and insightful they were.
Opus 5 first built me a detailed plan, but a couple important details were either obviously wrong or felt unnecessary. I went back and forth asking for sources and more information probably like 4 times and every time it did the "in looking at things in more detail it appears my previous advice was incorrect" spiel. It just became exhausting at some point because it feels like it really lays bare how LLMs are just minimizing that loss function but don't actually "understand" anything. It was really useful as a search engine (it correlated some highly relevant source docs), but I just couldn't trust it to believe it was actually done at any step.
I know it sounds silly but 1 layer ends uo being way worse than 2.
i can't even get agents to remember core instructions like "use jq instead of writing a python script to parse some json"..
I'm researching causal inference right now, and my main goal was to make sure I understand how to test estimation on synthetic data.
Basically, it's the same way it works with people. If you delegate a task that you don't understand, and you can't have a credibility proof (i.e. doctors, lawyers), then you research a topic well enough to be able to (1) define the task and (2) verify the end result.
One thing I wonder is, do you mentally 'fight back' monotonicity of your interactive tool? All seem to be in 3D space, with low-poly, like in a factory moving through the belt and giving you an information + textual description to read more
But sometimes you want to visualize the charts, or graph of simulations, or maybe even the parts of an item in the rocket.
In particular:
- I limit it/encourage it to give me single sentence questions
- I sometimes will ask it to tell me a motivating, human-grounded story, when we're starting a new concept: claude responds "Maya is a bond portfolio manager, and her boss has asked her to quickly price in what happened if yields go down. She knows her bond's average duration, a measure in time, but she doesn't have a percentage, which is what her manager wants. How can she give him a percentage number with just a duration figure and the proposed new yield?"
- I'll often ask claude to let me work through it, to derive the thing myself, often resulting in a string of thoughts with "yes/no" trailers, to get the LLM to reply yes or no only, and avoid derailing my train of thought. If yes, my train of thought keeps going. If no, I've got something wrong.
- I'll sometimes stop and have it craft an artifact. I typically say "build me a Brilliant.org-style interactive demo of the topic", especially when we get into the realm of looking at the actual maths of a thing (for which prose and dialog is not optimal by itself AFAICT)
- I'll do this while I'm traveling, while I'm walking, while I'm doing chores.
It's so much fun.
""" In this project, I require a socratically delivered line of conversation. Here's the typical structure to the conversation. I ask some question. You need to factor and reason about how to conduct and deliver a conversation. Best practices would be to limit terminology, or assess with the user whether they have a firm grasp on terminology before you use it. You must be very strict about this, it's unacceptable to just introduce a new concept, actor, phrase or other complication into the conversation without first labeling who what or why it exists for the conversation.
Conversation structure needs to be front-loaded with a brief interview for the user, "you understand X?", "whats your understanding of Y?".
Conversation structure then needs to proceed with single-sentence questions from the agent. User replies with an answer. Sometimes the agent needs to correct the user, but only ever do so with yet another question. """
^^ these are the instructions I have installed at the root of a "project".
Keep in mind, this is claude opus 5 low effort we're talking about, in the "projects" area of the mobile app. Here's the process I use to set up its knowledge:
1. I take screenshots of the textbook on my iPhone, and upload a chapter at a time.
2. I have it summarize the chapter into markdown by analyzing screenshots. You could probably achieve this simpler, if you just had the textbook in PDF.
3. I walk with my boy Clau-crates.
I've done this for a couple of weeks and haven't seen it revert back into its typical context-dumping behavior.
On second read, there's probably some clean up I could do. Thanks for making me pull it out and look at it. Things that could probably be improved:
- tell it to cross-check resources online to further ground itself
- use simple, short sentences (long sentences make the brain blur a bit)
- not be sycophantic (it seems like project mode has discarded with my root-level anti-sycophancy prompt)
I still learn new stuff, but I’m afraid it won’t have any value in a year or so.
For example, I’m pretty good at optimizing low level stuff, but right now you can just ask LLMs to do so and they are pretty good at it. They will profile the code and suggest reasonable options like 90% of the time.
Trust me when I say that in the hands of someone who doesn't have your experience, the LLMs would not be getting the results you get.
You might think what you're doing is trivial, it may be sessions that flow roughly, "Instrument this, okay this part is slow, profile this part, OK read the profile output and suggest a better approach".
But your experience will be steering it in the right direction, and you're probably unaware of just how much your experience is doing that guiding, as the LLM shoots off at 100mph, you feel like it's taking you with it, but you will be guiding it a lot more than you realise, and that's where learning and experience comes in, even if you're no longer operating at the lowest depth, your knowledge of that layer will be helping.
If nothing else, the experience to know when something is actually slow is a skill in itself. If a function takes 200ms, sometimes that's as quick as it can realistically go, and sometimes that's literally a million times slower than it could be, and there's actual skill and experience wrapped up in knowing what "slow" looks like.
And the craft is loose term, it can mean anything you like to get better at.
“asking the right questions” is also a moving target with each model release
People simply underestimate the value of doing the work and think that the end result is all that matters
https://en.wiktionary.org/wiki/eat_one%27s_seed_corn#English
What is fascinating is how you can witness it at so many levels of organization. One example: Employer executive get enamored with moving from labor to capital. They believe that by using LLMs, they can replace a lot of workers. At my place of employment, we have people that are surprised they can't file a Jira ticket describing a product ask, and have it kick off an implementation. You can build the skill to attempt that, but invariably you'll get back questions like "what do you mean by <x>" and "what do you want to do in this case, a, b, or c?"; questions that a product person or an exec are not well suited to answer.
In the past, programmers did that kind of interpretation and judgment call. So then you're in a quandary; who should do that work? Work that previously, you never imagined was an inherent part of what the replaceable code monkeys do at your beck and call?
And then, how do you hire for that? How do you find the training for the people that are experienced enough with... something... to know what a cohesive error response is, or what kind of telemetry strategy is best for that particular product and organization, what collection of product asks are incredibly complicated for what they're asking and can deliver 95% of the benefits at 5% of the work if we just do this instead, and whether you want to aim more towards thick or thin clients?
Who are those people? Wait, those are programmers? Wait, there's this whole collection of inherently human skills that we devalued, by not appreciating they were always quietly doing that for us in the past?
That's just one example. There's a repeating pattern of discovering where the work truly is, work that was embedded in manual patterns we might not have to involve ourselves with anymore, but is yet still essential. So the nature of our jobs changes massively, but the overall level of employment does not.
At least, not in the medium to long term. There is a lot of painful churn we have to suffer through first.
I wasn’t even really concerned with optimizing low level code before LLMs and that wasn’t why I was hired either.
However following that low level thread: We can look at the reasonable options and immediately know if they’re reasonable or nonsense. Why? We know the code. Now zoom a level out, where I think our expertise really lies.
Building a complex system isn’t easy. There are customers with requirements, there are budgets, SLAs etc. Sometimes one customer needs X and one needs Y. Our expertise is taking all of this in, and producing something that balances all the different variables. It’s knowing that we’ll expect X events a second so we’ll need Y to ensure we can tolerate failure.
Is it possible LLMs will be able to do all of that too? Maybe. But then why would our customers need the enterprises they pay for?
They tell you have "hit the nail on the head" when you really haven't.
They tell you have had a "great insight" when you are really haven't.
They give you the illusion of learning and progress but essentially give you faulty preconceptions will trip you up further down the road.
You can ask the LLM to be more critical and less sycophantic but that only gets you so far:
They want you to continue using, being dependent on and feeding data into the LLM--your independence isn't a priority.
I have stuff to do now, the value of the knowledge in a year or two isn't important if it solves the issues I have today.
I'm not 'wasting time' but I'm also not really learning.
The more things you understand, the higher the chance you'll spot a situation to use them in the future.
I think the best innovations come from times when someone is uniquely able to combine two of their previous experiences together. The more experiences you have in your back pocket the more combinations you have access to and the more likely you'll have a unique combination when the right problem comes along.
In particular it might be valuable to be in the habit of learning things that one is bad at doing.
Or not.
This is silly. This would be like arguing that encyclopedias made knowing things pointless. I learn new stuff for me.
Professionally, it's important to know enough to know if you're going in the correct direction. Practically, tokens are going to continue to cost money and knowledge can save you tokens.
1. It satisfies you curiosity (and curiosity is always valuable)
2. You can better utilize the LLM to expedite something you now have knowledge about
3. You still improve as an engineer/programmer/prompter/whatever
I still think it's very important not to outsource everything to AI because there is a lot of value in learning and doing things yourself which is an important part of life.
I can tell you that there is an enormous gap in ability between them despite them both using LLMs for daily IR work.
The reasons aren’t complicated. The senior responders have tacit knowledge of how breaches evolve and what to look for which gives them a much better framework for where to employ the LLM.
The juniors will normally start from “here are some logs, look for weird” which is fine but leads to tunnel vision and a lack of confidence in their reporting.
I don’t mandate that anyone do work with or without an LLM. I hire seniors based on experience and juniors based on interest. But my experience has so far been that our best up and comers focusing more on learning the technologies instead of leaving those details to the LLM are developing their intuition and understanding faster and in a more robust manner.
Makes sense.
> I ask it to review the accuracy of the knowledge base it built in the previous step.
Ooookay that sounds good.
> I proceed asking it to build a simulation of that topic in a low-poly, Rollercoaster Tycoon-like animation.
wat.
The main idea is you can do any style you want or like to learn.
In my experience it's infinitely easier and faster to learn deep, "boring" things when you understand how they relate to your shallow and wide understanding of all of the related components.
The LLM is merely a tool. And you can use it for domain discovery that enables efficient deep learning at an unprecedented rate or you can develop a cursory understanding of a topic and think yourself an expert.
It’s true of most things. Running, dieting, weightlifting being uncomfortable is a sign of progress.
This is the only path to mastery, or understanding if one prefers. There are no shortcuts to a person achieving deep understanding (a.k.a. "Aha!" moments).
Can a tool such as GenAI be beneficial to someone who already has done the work to understand? Absolutely. But it cannot infuse mastery into a person simply by its use.
Only the time and effort a person devotes can do that.
Another useful approach has been asking Codex to implement complex things, like a Kademlia DHT or BitTorrent client in a literate style with the explicit purpose to increase understanding by reviewing the source code.
Examples: https://rickcarlino.com/notes/note-dump-and-ai-summaries/ind...
https://github.com/RickCarlino/tiny-bt
That's actually a fun way to learn processes!
Totally agree, unfortunately careful simulation games are very rare
Otherwise you probably get more confused as you have mentioned.
On the other side, Peter Diamandis describes a situation where a bunch of kids were given a internet-connected computer and they had no teacher. Instead of it there was a “grandma” that checked kids from time to time.
After that there was a knowledge test that revealed “no teacher” approach was more efficient.
But it was a group, not an individual activity…
For example, “explain how the code in this file works,” I am familiar with the overall codebase, I know the purpose of the file, and I can read it or write tests to verify if I suspect what it’s telling me isn’t correct. Or, if it’s really important, I can overcome my introvertedness and ask the team member who wrote it…but that’s a last resort nowadays, which I am very thankful for. In 99% of cases since at least Claude 4.2 days, Claude and Codex have been very accurate. Gemini on the other hand messes up more frequently and sometimes does weird things like try to delete files it’s not familiar with, at least the 3.6 flash model I’ve been using lately does this. But, code explanations are still good for the most part.
You can ask the LLM how to do this. Start with a topic you know well to get the mechanism working and trust it well.
I assume this will become less of an issue in the future as there is more trust between the AI tools and me.
How does he know?
> Time and time again, when talking to people who rely on ChatGPT, Claude, Perplexity, and other general AI tools, I hear them say, “AI is incredible. It handles nearly everything I throw at them.”
> “What does it fumble with?” I’ll ask.
> “Well, it still gets things wrong when it comes to my line of work.”
https://www.dbreunig.com/2025/04/08/on-ai-observational-comi...
To be clear, I say "inaccurate" rather than "wrong" in this case because even if the information it returns is factually correct to the question being asked, students don't have an understanding of the complexity of the interdependent tectonic, regulatory, and spatial / experiential factors of a building sophisticated enough to ask their questions of the specificity and nuance necessary to get a good output that addresses the entire problem.
Anyway - with the students still learning to ask questions the right way, and the conditionally-incorrect facts making their learning more complicated rather than less, I hit on a strategy for them to use LLM's that seemed to help much better.
I suggested that instead of ask the LLM for the factual answer, or even better for the facts and an explanation, that they ask it to direct them to the proper place in the source material to find the answer themselves. Then, to treat it like a lab partner. IE:
Hey Claude I'm looking for "x."
Claude: "look at foo, bar."
Thank you - chapter (foo) part (bar) table (goo) says "car." However I notice that footnote (hoo) says there's an exception if "dar." Which is what I have. Walk me through this exception...
It seemed to have good results as a guide to understanding the disparate bodies of knowledge that they will eventually have to keep together in their heads and work synthetically and non-linearly through, rather than just as an external source of blindly trusted authority.
Hallucinations bring to question what you think you've learned. That's going to cost long-term if you labor under mis-apprehensions until you maybe figure out you learned something wrong.
Yes, sometimes its wrong, most times its right, cross checking is fairly easy, not using it because of the possibility its wrong seems a baby with the bath water thing.
I learnt a lot of functional programming from it, stuff I've always wanted to learn, but just didn't have the time and really the sources can be difficult, it really explained things well, and as someone else said in this thread, you can ask questions over and over until you understand, asking a person that (if you can get an expert) would drive them nuts. Maybe my experience isn't typical, its hard to tell, everyone reports something different.
having things defined/have correct solution to compare for review is useful, and keeps llm on track. Don't think i would trust llm if it were reviewing it all on its own
Colleagues often suggest podcasts and videos - I very, very rarely listen to them or see them.
The bandwidth is too low. It's not efficient and ultimately I'm bored.
This is a nice project, it looks cute. I watched some of the pages But I want more than that, more information, and faster - still a Wiki fan.
Also, step number 2 in the flow: have the LLM check itself... Naah, I don't believe that.
But you're not the only using gen ai like that. Take care.
Indeed it's often a waste of time to just focus on talking people fully if you want to learn fast, reading and especially deliberate practice are better for that. But if you don't have the time, energy or focus, then listening to interviews in the background can be useful supplementally
OTOH, I am curious if there's a "practical value" to this exercise? If the LLM already contains the information and implementation knowledge to implement the networking stack inside an FPGA by itself, what value do I gain by learning about HDL, TCP, the bespoke Xillinx tooling, reading the documentation, reading papers on the implementation and going through every bit of details and theory. I feel like there's a meta skill that is more worthwhile for "practical value".
The skill then riffs with me, judging my ideas and suggesting alternatives. We go back and forth until something useful comes out of it. This process isn’t unlike how I do normal development.
However, once agreed it breaks the work into “steps”. It then creates a tutorial for me, for those steps, explaining each line, why each change happens etc. I can then ask questions, muse about an alternative idea etc. Then I do the steps, and I’ve learned and gotten what I wanted to get done.
This has been how I’ve been learning Godot and making a game for the past month or so. I didn’t go in blind, I started with a course from GDQuest so I could feel confident guiding the tutorials. I will say though, having a tutor to bounce ideas off of has been really useful.
I still try to figure it out myself, consult the docs, discord etc. But if I’m stumped I’ll run my tutor skill and have some fun.
I have done this for all my work this week and it works quite well.
For one it lets you actually query the LLM as to why, their plans give a high level not every single change and it allows you to correct it as you go and the plan will change.
As for how useful it is to understand thins, I believe it's still useful and hope it will continue to be.
Cat also has an awesome podcast with her wife, Ashley Juavinett, Phd, called Change, Technically.[4]
I encourage everyone to check out her work! She’s dedicated her life to helping software developers get the support they need inside organizations to be seen as humans, not just robots.
1: https://www.drcathicks.com#book 2: https://github.com/DrCatHicks/learning-opportunities 3: https://github.com/DrCatHicks/learning-goal 4: https://www.changetechnically.fyi
Does anyone else use Claude like this?
It's sped up my learning by 10x. I struggled with 'just reading a book.' Take kubernetes. I hemmed and hawed and spent years periodically reading some dry book or blog or official doc, falling asleep, and forgetting while I got busy. Now I'm aggressively working with it, almost like I'm addicted to a gamification, of getting through our learning timeline, and I'm excited to move forward as quickly as possible and pass its tests.
It's like a fake teacher, because I can also ask it to drill into a topic or re-explain itself if it made no sense.
The only thing that worries me is, sometimes I'll say something like, "Um, are you sure about that?", and it'll apologize and correct itself. I barely challenged it!
I'm using LLMs right now to build a terminal browser, a GUI browser, and a PyTorch/LibTorch replacement. It's really fun to be able to learn and make progress this way. It's like reading multiple interactive books, where every concept can be explained again and again until I understand it.
I had a similar realization a few months back and am working on a tool that generates "mermaid walkthroughs". It is 1000% less pretty but it is fast and is pretty good at explaining how services work or what a code review does or just as a way for your agent to explain some decision to you.
https://github.com/scottrogowski/ariel
I don't think we are even that far when the complexity AI can handle surpasses 99.999% of what humans can handle, where AI make e.g. physics discoveries beyond the grasp of most humans and it will have to "dumb it down" when talking with humans -even physicists- but not with other AIs
LLMs and systems intersection - https://kernelspace.naigap.com
Distributed systems - https://byzantine.play.naigap.com
What is your process for creating these resources?
I find LLM's do great for learning when I ask what are the principles, how the main applications work, what are the key drawbacks, where are the growth plates in the field, etc. - the kind of thing a good advisor points to. Sometimes I have to ask it explicitly to use topological order of topics and show relations, which often highlights the gradient changes in the learning curve. For pruning, it's surprisingly good applying philosophical heuristics - Occam's razor, or Derrida's differance (the difference that makes a difference), etc.
And finally, no learning is effective without problem sets, and for those LLM's at times get me over blocking issues.
The degenerate case is memorizing the glib phrases regurgitated back to me; they're helpful and functional enough to get me into real trouble!
> What you get is a beautiful animation that is 100% accurate and free of hallucinations.
How do you make that leap?
I think that would be a way more natural way to explore than being stuck on the classic linear output of a LLM.
But
> What you get is a beautiful animation that is 100% accurate and free of hallucinations
100% free of hallucinations when you're not an expert that can check it is impossible. LLM hallucinations are an unsolved problem.
It's not perfect, but I'm optimistic this will be a useful way to teach/learn in the future.
And, to be clear, I think this will be best utilized within a group/community setting. I don't think it will replace teachers or classrooms.
Also worth mentioning that Matt Pocock has a /teach skill that creates interactive, learning sites for learning a new skill.
> [I ask it to build an interactive thing]
> I then push it to a new repo and enable GitHub Pages for it.
Congratulations. You are an echo chamber for LLMs. Use it to create, and verify, and post to then be scraped and trained on again.
I've been working on a similar process of pushing to github pages, but focused more on having "practice sessions" with coding blocks to test content. Using webassembly and mock servers to mock backend endpoints Here's one I built to build a full stack llm chat system in the browser.
https://model-systems-labs.github.io/latent/llm-systems/less...
How do you know if you're learning this for the first time? Very risky to learn from LLMs. I've done it, but you have to keep your wits about you. Lots of "oh of course you're right - what I just told you was completely wrong".
Last month, I read The Prince and had it make a text adventure campaign for me.
For a lot of other topics, I often just ask it to create a simple python example that I can run.
It reduces friction a ton, but at the end of the day I’m not skipping anything.
This is my biggest societal issue with LLMs: they allow you to think that you’ve “learned” a topic because you read a lot of technical terms. I don’t think it bodes well for the future.
0: https://en.wikipedia.org/wiki/Semiconductor_device_fabricati...
Inthink this was always the best way to learn. But it used to require immense work for a teacher.
My compiler course is a great example - program in plug in a stage of a compiler.
These exercises can be made on demand and incredibly easy now.
God does everything have to be productized and glorified as if you’ve invented a new way of learning. Read some books!
Surprisingly effective.
(Yes I confirmed it was ok to use AI with the material)
You're just memorizing a nonsensical recipe. What are the constraints? Why do we do X rather than Y? How does a particular thing scale? etc.
All you're doing is fooling yourself into thinking that you've acquired some knowledge. When in reality you haven't even learned the basic mental model to reason about this stuff.
You've learned something when you have a mental model that makes correct predictions. Until then you've memorized it at best, and as with most memorized things it will decay exponentially and will be gone from your memory soon enough.
Then I went back to a book written by humans (Eric Nikitin's "Realm of Oberon").
I also struggle with LLMs explaining things, but for the opposite reason.
I consistently have problems to get short, precise but plain/simple answers.
Instead I'm overwhelmed with walls of texts, often filled with jargon that is a mixture of imprecise and unneeded.
The style at which I learn better is by asking about stuff interactively. I ask you what something is, you give me a 3-4 sentences top answer. Then I explore and dig into the topic from your answer on the things I want to know better.
Do you know it is free of hallucinations because you crossed checked it with the source material or because you told the LLM "don't hallucinate"
Perhaps basic special relativity could be done this way, or simple derivatives, but definitely not general relativity or integrals, much less PDEs. I guess I was hoping for a 3B1B type output. Oh well.