Academy Speaking guide

CELPIP Speaking Task 3: describing a scene template

A CELPIP Speaking Task 3 template for describing a picture in 60 seconds, two worked answers, and the one step only about 14% of learners actually take.

HelloCelpip Team
16 min read

Use this CELPIP Speaking Task 3 template: open with one sentence saying where the scene is, then walk through it in a fixed order, giving each person a position, an appearance, and an action. Task 3 is Describing a Scene. You get 30 seconds to prepare and 60 seconds to speak, once, with no second take.

The instruction on screen contains the whole trick, and almost everyone reads past it: the person you are speaking to cannot see the picture. This guide gives you the template, two worked answers, and the one step our data shows almost nobody takes.

Task 3 at a glance
Preparation time
30 seconds, with the picture already on screen.
Speaking time
60 seconds, the standard length for six of the eight Speaking tasks.
What you must do
Describe what is happening in an illustration to someone who cannot see it.
What is on screen
A single drawn scene, usually a busy everyday place with several people doing different things.
Where it sits
Third of the eight Speaking tasks. Speaking is the last section of the test and runs about 15 minutes.

#What does CELPIP Speaking Task 3 ask you to do?

Task 3 shows you one illustration and asks you to describe some of what is happening in it. The scenes are ordinary and busy: a supermarket, a park, a waiting room, a street corner. Several people are usually doing several different things at once.

The instruction adds one condition that changes everything about how you should answer. Your listener cannot see the picture. You are not pointing at a scene you both share. You are building it in their head from nothing.

CELPIP Speaking Task 3 on HelloCelpip: an illustrated supermarket checkout scene where a bag of apples has spilled across the floor, shown beside the instruction to describe what is happening to someone who cannot see the picture, and a preparation countdown.
Speaking Task 3 (Describing a Scene) in HelloCelpip, during the 30 seconds of preparation.

Look at how much is going on in a typical scene. There is no way to cover all of it in 60 seconds, and you are not expected to. Paragon's own strategies say directly that you do not need to describe everything. Choosing what to leave out is part of the task.

Try a Speaking Task 3 question

#Which part of a Task 3 answer do learners actually miss?

This is the widest gap we have measured on any CELPIP task, in any skill.

We looked at 15130 anonymised, aggregated Speaking Task 3 responses on HelloCelpip and checked each one against what the task asks for. Our evaluator judged that about 93% described what is happening in the picture. Only about 14% painted a picture the listener could actually see.

Painted a picture the listener could see 14% Described what is happening in the picture 93%
What Speaking Task 3 responses on HelloCelpip actually did, across 15130 anonymised, aggregated responses. Percentages are our evaluator's assessment.
What the task asks for Share of responses that did it
Painted a picture the listener could see about 14%
Described what is happening in the picture about 93%

Nearly everyone does the literal task. They report the events accurately: a woman is standing, a boy is running, a man is paying. Almost nobody does the actual job, which is to give the listener enough that they could sketch the scene afterwards.

The requirement is published, as it was for the other tasks. Paragon lists "build a picture" and "describe the people's appearance, actions, and feelings" among its Task 3 strategies. And the instruction on screen states the reason out loud: your listener cannot see it.

#What is the difference between describing and painting?

Three things, and they are all easy to add. A reported event tells your listener what is happening. A painted scene also tells them where it is happening, who it is happening to, and how it seems to feel.

Reported Painted
A woman is picking something up. On the left, near the till, a woman in a green cardigan is crouching down to gather apples that have rolled across the floor.
A boy is standing there. Just in front of her, a small boy in a yellow hoodie has stopped mid-step, holding one apple and looking delighted with himself.
The store is busy. Behind them the queue has stopped moving, and a cashier in a blue apron is leaning over her screen with one hand raised, trying to get someone's attention.

The painted column is not more advanced English. Every word in it is ordinary. What it adds is a position, a colour or a piece of clothing, and a hint of how the person seems.

Exam Tip
Depth beats coverage
Paragon's strategies say to start with a general statement, then focus on some details, and that you do not need to describe everything. Three or four people described properly will score better than eight people listed. If you find yourself saying and there is also, you have switched from painting to listing.

#A CELPIP Speaking Task 3 template you can adapt

Template
A describing a scene answer you can adapt to any picture
This looks like [the place], and it seems to be [busy, quiet, early morning, the middle of an event].

In the foreground, [person 1: what they are wearing, where exactly, what they are doing]. [How they seem to feel.]

Just behind them, [person 2: same three things].

On the [left or right], [person 3, and what they are doing about it].

Overall, the scene feels [the general mood], mainly because [the one detail that creates it].

The words in brackets are instructions to you, not lines to copy. The template is built around a fixed movement through the picture, foreground first, then behind, then to one side, because a listener who cannot see the scene needs a route through it. Jumping between corners is what makes a description hard to follow, and that lands under Listenability.

Sentence starters that fit almost any picture:

  • The opening statement: "This looks like a…", "The scene takes place in…", "It seems to be a busy…"
  • Placing someone: "In the foreground…", "Just behind her…", "On the left of the picture…", "Off to one side…"
  • Describing them: "a man in a dark blue coat…", "an older woman wearing an apron…", "a boy of about seven…"
  • Their action: "is reaching for…", "has just knocked over…", "is leaning across the counter to…"
  • Their feeling: "he looks a bit embarrassed…", "she seems concerned…", "they are clearly enjoying it…"
  • Closing: "Overall it feels…", "The whole scene looks…"

Do not memorise a finished answer. Memorise the route through the picture, and let the scene supply the content.

#What are the 30 seconds of preparation for?

Choosing three or four people and deciding your route. The picture is already on screen during preparation, so unlike other tasks you are not thinking up content, you are selecting it.

How to spend the 30 seconds
  1. Name the place in your head
    Supermarket, park, clinic waiting room. This becomes your opening sentence and gives the listener somewhere to put everything else.
  2. Find the one thing going wrong
    Most scenes have a small incident at their centre: something spilled, someone late, a queue stuck. Build around it rather than starting at a random corner.
  3. Pick three or four people, not eight
    Choose the ones nearest the incident. Note a colour or a piece of clothing for each, because that is what makes them distinct to someone who cannot see them.
  4. Decide your route
    Foreground, then behind, then one side. Fixing the order now stops you jumping around when the recording starts.

Those thirty seconds are a separate skill from speaking, and you can train them on their own. Idea Coach walks you through an answer one piece at a time, so the shape is automatic before you ever start the clock.

Build your description for Task 3

#What does a weak Task 3 answer sound like?

It sounds accurate and completely invisible. Read this one and try to picture the scene.

HelloCelpip example
Answer 1: everything is true, nothing can be seen

The picture: an outdoor farmers market on a Saturday morning. You cannot see it, which is exactly the point.

In this picture I can see a market. There are many people and they are buying things. One woman is standing near a table and she is looking at some vegetables. There is a man who is selling fruit and he is talking to a customer. Some children are playing nearby. There is also a dog. The weather looks nice and everyone seems happy. It is a busy day at the market and people are enjoying their time there.

Why it falls short: Every sentence is grammatically fine and factually true, and it would be judged as having described what is happening. But nothing has a position, nobody has an appearance, and no two sentences connect. You could not draw this. It is also a list: seven separate people and things, none of them developed. At about 79 words it leaves roughly a third of the minute empty.

This is what nearly nine in ten Task 3 answers look like. Notice that fixing it does not require better English.

#What does a strong Task 3 answer sound like?

The same market, the same ordinary vocabulary, with a place, a route, and three people you can actually see.

HelloCelpip example
Answer 2: the same market, painted

The picture: the same outdoor farmers market on a Saturday morning.

This looks like an outdoor farmers market on a bright Saturday morning, and it is clearly the busiest part of the day.

In the foreground, a woman in a long red coat is leaning over a table of vegetables with a tomato in each hand, comparing them. She has a canvas bag hooked over one elbow that is already full, so she has been there a while.

Just behind her, the stallholder is a heavy-set man in a green apron, holding up a paper bag of peaches and laughing at something his customer has said. Neither of them is in any hurry.

Off to the left, two small children have crouched down to pet a large brown dog that is tied to the corner of a stall, while the dog patiently ignores them.

Overall the scene feels relaxed and slightly crowded, mostly because nobody in it is moving quickly.

Why it works: The first sentence puts the listener somewhere. Then the answer moves in a straight line, foreground, behind, left, so the scene assembles in order instead of jumping around. Three people get a position, a piece of clothing, an action, and a hint of feeling, and one small detail does extra work: the bag already full tells you she has been shopping a while. At 149 words it fills the full minute, a little above the 120-word average, at a pace that is comfortable rather than hurried.

Read the two out loud one after the other. The second is not harder English, and it does not cover more of the market. It covers less of it, in more depth.

#How much do you need to say in 60 seconds?

About 120 words, at an ordinary pace. That is roughly three sentences for each person you describe, plus an opening and a closing line.

Across Task 3 responses recorded on HelloCelpip, the average answer runs about 120 words and uses almost the full minute. That figure comes from our current evaluator, which records exactly what was said, including natural pauses and fillers.

Three or four people at three sentences each is what fills a minute comfortably. This is the arithmetic reason depth beats coverage: eight people cannot get three sentences each, so listing them is what leaves an answer sounding thin even when it is the full length.

#What do CELPIP raters listen for?

Speaking responses are rated in four areas, and the official CELPIP Performance Standards give a plain question for each one.

The four rated areas
Content / Coherence
How well are your ideas organized and developed? A fixed route through the picture is what organisation means here, and detail is what development means.
Vocabulary
What is the range of your vocabulary and can you use it naturally? Descriptive everyday words, especially for position, colour, and manner.
Listenability
How easy is it to listen to and understand your response? Jumping around the picture is what most often makes a description hard to follow.
Task Fulfillment
How well did you follow the instructions and use an appropriate tone? Describe the scene to someone who cannot see it, and keep going for the full time.

There is a pattern in where marks go. Across more than 13,500 anonymised, aggregated Task 3 responses on HelloCelpip, about a third of everything our evaluator flags is a vocabulary problem, more than any other area. The other three sit close together behind it, each accounting for roughly a fifth to a quarter of what gets flagged.

On this task specifically, the vocabulary that pays is not advanced. It is the ordinary descriptive words most learners already know and do not use: positions such as beside, behind, and in the corner, and manner words such as carefully, patiently, and awkwardly.

#Common Task 3 mistakes

What costs marks in a describing a scene answer
  • Listing everything in the picture instead of describing a few things properly. This is the biggest gap in our data by a wide margin.
  • Starting with I can see and repeating it for every sentence.
  • Giving no positions, so the listener cannot assemble the scene.
  • Jumping between corners of the picture with no route.
  • Guessing at a backstory instead of describing what is actually shown.
  • Forgetting that the listener cannot see it, and saying things like this one here.
  • Running out at 35 seconds because everything was listed rather than developed.

If you feel yourself drying up, do not look for another person to add. Go back to someone you have already mentioned and say what they are wearing, or how they seem to feel about what is happening.

#How to practise Task 3

Practice
Record one describing a scene answer today
Speak for the full 60 seconds on a real Task 3 picture, then get a report across the four rated areas with the wording that cost you marks quoted back from your own answer.
Six of the eight Speaking tasks give you 60 seconds, so the pacing you build here carries across.

There is one practice exercise worth more than the rest on this task. Record your answer, then play it back to someone who has not seen the picture and ask them to describe what they imagined. Whatever they get wrong is the detail you left out. If you are practising alone, write down what your own description would let a stranger draw.

A simple way to use it: record two or three Speaking Task 3 questions, then take a full mock test when you want to hear yourself under real timing.

If you are working through the Speaking tasks in order, the CELPIP Speaking Task 1 template covers Giving Advice, which is the one task that gives you 90 seconds instead of 60.

#Frequently asked questions

Frequently Asked Questions
What is CELPIP Speaking Task 3?
Task 3 is called Describing a Scene. You are shown one illustration of an everyday place, usually with several people doing different things, and asked to describe what is happening in it to someone who cannot see the picture.
How long do you get for CELPIP Speaking Task 3?
Paragon’s CELPIP Speaking Pro study pack lists 30 seconds of preparation and 60 seconds of speaking for Task 3. The picture is on screen during the preparation time, and recording starts automatically when the clock ends.
Do I have to describe everything in the picture?
No, and trying to is the most common mistake. Paragon’s strategies say to start with a general statement, then focus on some details, and state directly that you do not need to describe everything. Three or four people described properly fills the 60 seconds better than a list of eight.
What do learners get wrong most often in Task 3?
Describing the events without making the scene visible. Across 15130 anonymised, aggregated Task 3 responses on HelloCelpip, our evaluator judged that about 93% described what is happening but only about 14% painted a picture the listener could see. That is the widest gap we have measured on any CELPIP task.
How do I make a description easier to follow?
Move through the picture in a fixed order and say where each person is. Foreground, then behind, then one side. A listener who cannot see the scene is assembling it as you speak, so a route makes it far easier to follow, and jumping between corners is scored under Listenability.
What tense should I use in Speaking Task 3?
The present, and usually the present continuous, because you are describing what is happening now: she is reaching for, he is leaning over, they are waiting. This is the opposite of Task 2, which is a past story.
How many words should a Task 3 answer be?
There is no word requirement, since it is a spoken task. As a guide, Task 3 answers on HelloCelpip average about 120 words across the 60 seconds, which is an ordinary, unhurried pace. That works out at roughly three sentences for each of three or four people.
Official sources

Confirm the eight Speaking tasks and the 15-minute Speaking section on the official CELPIP Test Format page. Per-task timing of 30 seconds to prepare and 60 seconds to speak, and the Task 3 strategies quoted above, are published in Paragon’s CELPIP Speaking Pro study pack, and the four rated areas come from the CELPIP Speaking Performance Standards. Your HelloCelpip level uses the same 0 to 12 CELPIP scale, so you can track real progress between now and test day. Your official score comes from Paragon on test day. HelloCelpip is an independent study resource and is not affiliated with CELPIP or Paragon Testing Enterprises.

Official sources:

Last updated: July 24, 2026