Module
Module 4 of 5Lesson 1 of 1~14 min

Reasoning and reading images, what these capabilities really change

Lesson 1 of the module "Reasoning and multimodality" in the course "Understand what LLMs do well, and where they fail".

Lesson objective

By the end of this lesson, you will be able to decide when to turn on an assistant's reasoning mode, and to spot in the reading of an image, screenshot or PDF the elements that need checking.

Where it fits

Reasoning and multimodality

What changes when a model "thinks" before answering or reads images, and what does not change?

Lessons in this module

  1. Reasoning and reading images, what these capabilities really change (this lesson)

What you will learn in the course

This lesson is part of the course Understand what LLMs do well, and where they fail

  • Explain how an LLM produces an answer (tokens, next-token prediction, pretraining then post-training) and why a fluent answer is not a verified answer.
  • Anticipate hallucinations, sycophancy, poorly calibrated confidence and run-to-run variability, and choose the right check for each case.
  • Locate what the model knows (training data, knowledge cutoff, uneven coverage) and what web search, provided documents and connectors change.
  • Explain the context window as working memory (limit, degradation over long contexts, no memory across conversations without a dedicated feature) and derive working rules from it.
  • Assess what reasoning (thinking) and reading images or documents bring, and what they do not guarantee.
  • Sort the tasks of your job into "delegate / verify / keep" based on the cost of an error, how easy it is to verify, the model property involved and the value of doing it yourself.