Use this CELPIP Speaking Task 3 template: open with one sentence saying where the scene is, then walk through it in a fixed order, giving each person a position, an appearance, and an action. Task 3 is Describing a Scene. You get 30 seconds to prepare and 60 seconds to speak, once, with no second take.
The instruction on screen contains the whole trick, and almost everyone reads past it: the person you are speaking to cannot see the picture. This guide gives you the template, two worked answers, and the one step our data shows almost nobody takes.
#What does CELPIP Speaking Task 3 ask you to do?
Task 3 shows you one illustration and asks you to describe some of what is happening in it. The scenes are ordinary and busy: a supermarket, a park, a waiting room, a street corner. Several people are usually doing several different things at once.
The instruction adds one condition that changes everything about how you should answer. Your listener cannot see the picture. You are not pointing at a scene you both share. You are building it in their head from nothing.
Look at how much is going on in a typical scene. There is no way to cover all of it in 60 seconds, and you are not expected to. Paragon's own strategies say directly that you do not need to describe everything. Choosing what to leave out is part of the task.
Try a Speaking Task 3 question →
#Which part of a Task 3 answer do learners actually miss?
This is the widest gap we have measured on any CELPIP task, in any skill.
We looked at 15130 anonymised, aggregated Speaking Task 3 responses on HelloCelpip and checked each one against what the task asks for. Our evaluator judged that about 93% described what is happening in the picture. Only about 14% painted a picture the listener could actually see.
| What the task asks for | Share of responses that did it |
|---|---|
| Painted a picture the listener could see | about 14% |
| Described what is happening in the picture | about 93% |
Nearly everyone does the literal task. They report the events accurately: a woman is standing, a boy is running, a man is paying. Almost nobody does the actual job, which is to give the listener enough that they could sketch the scene afterwards.
The requirement is published, as it was for the other tasks. Paragon lists "build a picture" and "describe the people's appearance, actions, and feelings" among its Task 3 strategies. And the instruction on screen states the reason out loud: your listener cannot see it.
#What is the difference between describing and painting?
Three things, and they are all easy to add. A reported event tells your listener what is happening. A painted scene also tells them where it is happening, who it is happening to, and how it seems to feel.
| Reported | Painted |
|---|---|
| A woman is picking something up. | On the left, near the till, a woman in a green cardigan is crouching down to gather apples that have rolled across the floor. |
| A boy is standing there. | Just in front of her, a small boy in a yellow hoodie has stopped mid-step, holding one apple and looking delighted with himself. |
| The store is busy. | Behind them the queue has stopped moving, and a cashier in a blue apron is leaning over her screen with one hand raised, trying to get someone's attention. |
The painted column is not more advanced English. Every word in it is ordinary. What it adds is a position, a colour or a piece of clothing, and a hint of how the person seems.
#A CELPIP Speaking Task 3 template you can adapt
This looks like [the place], and it seems to be [busy, quiet, early morning, the middle of an event].
In the foreground, [person 1: what they are wearing, where exactly, what they are doing]. [How they seem to feel.]
Just behind them, [person 2: same three things].
On the [left or right], [person 3, and what they are doing about it].
Overall, the scene feels [the general mood], mainly because [the one detail that creates it].
The words in brackets are instructions to you, not lines to copy. The template is built around a fixed movement through the picture, foreground first, then behind, then to one side, because a listener who cannot see the scene needs a route through it. Jumping between corners is what makes a description hard to follow, and that lands under Listenability.
Sentence starters that fit almost any picture:
- The opening statement: "This looks like a…", "The scene takes place in…", "It seems to be a busy…"
- Placing someone: "In the foreground…", "Just behind her…", "On the left of the picture…", "Off to one side…"
- Describing them: "a man in a dark blue coat…", "an older woman wearing an apron…", "a boy of about seven…"
- Their action: "is reaching for…", "has just knocked over…", "is leaning across the counter to…"
- Their feeling: "he looks a bit embarrassed…", "she seems concerned…", "they are clearly enjoying it…"
- Closing: "Overall it feels…", "The whole scene looks…"
Do not memorise a finished answer. Memorise the route through the picture, and let the scene supply the content.
#What are the 30 seconds of preparation for?
Choosing three or four people and deciding your route. The picture is already on screen during preparation, so unlike other tasks you are not thinking up content, you are selecting it.
-
Name the place in your headSupermarket, park, clinic waiting room. This becomes your opening sentence and gives the listener somewhere to put everything else.
-
Find the one thing going wrongMost scenes have a small incident at their centre: something spilled, someone late, a queue stuck. Build around it rather than starting at a random corner.
-
Pick three or four people, not eightChoose the ones nearest the incident. Note a colour or a piece of clothing for each, because that is what makes them distinct to someone who cannot see them.
-
Decide your routeForeground, then behind, then one side. Fixing the order now stops you jumping around when the recording starts.
Those thirty seconds are a separate skill from speaking, and you can train them on their own. Idea Coach walks you through an answer one piece at a time, so the shape is automatic before you ever start the clock.
Build your description for Task 3 →
#What does a weak Task 3 answer sound like?
It sounds accurate and completely invisible. Read this one and try to picture the scene.
The picture: an outdoor farmers market on a Saturday morning. You cannot see it, which is exactly the point.
In this picture I can see a market. There are many people and they are buying things. One woman is standing near a table and she is looking at some vegetables. There is a man who is selling fruit and he is talking to a customer. Some children are playing nearby. There is also a dog. The weather looks nice and everyone seems happy. It is a busy day at the market and people are enjoying their time there.
Why it falls short: Every sentence is grammatically fine and factually true, and it would be judged as having described what is happening. But nothing has a position, nobody has an appearance, and no two sentences connect. You could not draw this. It is also a list: seven separate people and things, none of them developed. At about 79 words it leaves roughly a third of the minute empty.
This is what nearly nine in ten Task 3 answers look like. Notice that fixing it does not require better English.
#What does a strong Task 3 answer sound like?
The same market, the same ordinary vocabulary, with a place, a route, and three people you can actually see.
The picture: the same outdoor farmers market on a Saturday morning.
This looks like an outdoor farmers market on a bright Saturday morning, and it is clearly the busiest part of the day.
In the foreground, a woman in a long red coat is leaning over a table of vegetables with a tomato in each hand, comparing them. She has a canvas bag hooked over one elbow that is already full, so she has been there a while.
Just behind her, the stallholder is a heavy-set man in a green apron, holding up a paper bag of peaches and laughing at something his customer has said. Neither of them is in any hurry.
Off to the left, two small children have crouched down to pet a large brown dog that is tied to the corner of a stall, while the dog patiently ignores them.
Overall the scene feels relaxed and slightly crowded, mostly because nobody in it is moving quickly.
Why it works: The first sentence puts the listener somewhere. Then the answer moves in a straight line, foreground, behind, left, so the scene assembles in order instead of jumping around. Three people get a position, a piece of clothing, an action, and a hint of feeling, and one small detail does extra work: the bag already full tells you she has been shopping a while. At 149 words it fills the full minute, a little above the 120-word average, at a pace that is comfortable rather than hurried.
Read the two out loud one after the other. The second is not harder English, and it does not cover more of the market. It covers less of it, in more depth.
#How much do you need to say in 60 seconds?
About 120 words, at an ordinary pace. That is roughly three sentences for each person you describe, plus an opening and a closing line.
Across Task 3 responses recorded on HelloCelpip, the average answer runs about 120 words and uses almost the full minute. That figure comes from our current evaluator, which records exactly what was said, including natural pauses and fillers.
Three or four people at three sentences each is what fills a minute comfortably. This is the arithmetic reason depth beats coverage: eight people cannot get three sentences each, so listing them is what leaves an answer sounding thin even when it is the full length.
#What do CELPIP raters listen for?
Speaking responses are rated in four areas, and the official CELPIP Performance Standards give a plain question for each one.
There is a pattern in where marks go. Across more than 13,500 anonymised, aggregated Task 3 responses on HelloCelpip, about a third of everything our evaluator flags is a vocabulary problem, more than any other area. The other three sit close together behind it, each accounting for roughly a fifth to a quarter of what gets flagged.
On this task specifically, the vocabulary that pays is not advanced. It is the ordinary descriptive words most learners already know and do not use: positions such as beside, behind, and in the corner, and manner words such as carefully, patiently, and awkwardly.
#Common Task 3 mistakes
- Listing everything in the picture instead of describing a few things properly. This is the biggest gap in our data by a wide margin.
- Starting with I can see and repeating it for every sentence.
- Giving no positions, so the listener cannot assemble the scene.
- Jumping between corners of the picture with no route.
- Guessing at a backstory instead of describing what is actually shown.
- Forgetting that the listener cannot see it, and saying things like this one here.
- Running out at 35 seconds because everything was listed rather than developed.
If you feel yourself drying up, do not look for another person to add. Go back to someone you have already mentioned and say what they are wearing, or how they seem to feel about what is happening.
#How to practise Task 3
There is one practice exercise worth more than the rest on this task. Record your answer, then play it back to someone who has not seen the picture and ask them to describe what they imagined. Whatever they get wrong is the detail you left out. If you are practising alone, write down what your own description would let a stranger draw.
A simple way to use it: record two or three Speaking Task 3 questions, then take a full mock test when you want to hear yourself under real timing.
If you are working through the Speaking tasks in order, the CELPIP Speaking Task 1 template covers Giving Advice, which is the one task that gives you 90 seconds instead of 60.
#Frequently asked questions
What is CELPIP Speaking Task 3?
How long do you get for CELPIP Speaking Task 3?
Do I have to describe everything in the picture?
What do learners get wrong most often in Task 3?
How do I make a description easier to follow?
What tense should I use in Speaking Task 3?
How many words should a Task 3 answer be?
Confirm the eight Speaking tasks and the 15-minute Speaking section on the official CELPIP Test Format page. Per-task timing of 30 seconds to prepare and 60 seconds to speak, and the Task 3 strategies quoted above, are published in Paragon’s CELPIP Speaking Pro study pack, and the four rated areas come from the CELPIP Speaking Performance Standards. Your HelloCelpip level uses the same 0 to 12 CELPIP scale, so you can track real progress between now and test day. Your official score comes from Paragon on test day. HelloCelpip is an independent study resource and is not affiliated with CELPIP or Paragon Testing Enterprises.
Official sources:
Last updated: July 24, 2026