Use this CELPIP Speaking Task 4 template: name the moment in the picture that is about to change, predict what happens next, then say what in the image tells you so. Repeat for a second prediction. Task 4 is Making Predictions. You get 30 seconds to prepare and 60 seconds to speak, once, with no second take.
Task 4 comes straight after Task 3 and shows you the same picture again. The question changes from what is happening to what will most probably happen next. This guide gives you the template, two worked answers, and the habit that separates a prediction from a guess.
#What does CELPIP Speaking Task 4 ask you to do?
Task 4 asks what you think will most probably happen next in the picture. The wording matters: most probably. You are not being asked to invent the most interesting ending. You are being asked to read the scene and say where it is heading.
The picture is the one you have just spent 60 seconds describing in Task 3. That is a real advantage, and most people waste it. You already know where everyone is standing and what they are doing, so the 30 seconds of preparation can go entirely on the future rather than on re-reading the image.
Notice how much of the prediction is already sitting in the image. A small boy stretching up towards a display case, a tray of full coffee cups being carried through a crowd, a queue backed up to the door. Good scenes are built with several things mid-motion, and those are your answer.
Try a Speaking Task 4 question →
#Which part of a Task 4 answer do learners actually miss?
Task 4 is one of the two Speaking tasks our data shows learners handle well. There is no large failure here, which is itself worth knowing, and there is one modest gap worth closing.
We looked at 13960 anonymised, aggregated Speaking Task 4 responses on HelloCelpip and checked each one against what the task asks for. Our evaluator judged that about 95% made predictions rather than describing the scene, and about 93% made more than one prediction. About 87% tied their predictions to something actually visible in the picture.
| What the task asks for | Share of responses that did it |
|---|---|
| Tied the predictions to the picture | about 87% |
| Made more than one prediction | about 93% |
| Made predictions rather than describing | about 95% |
Roughly one answer in eight drifts away from the image. The prediction is reasonable, the English is fine, and nothing in the scene supports it: the customer will complain to the manager, the shop will close early, the two of them will become friends. Those are stories, not predictions.
That matters because the marks for Task 4 are not awarded for imagination. They are awarded for reasoning that a listener can follow, and a listener can only follow reasoning that starts from something they can see.
#What turns a guess into a prediction?
One clause. Say what you think will happen, then say what in the picture tells you so.
| Guess | Prediction |
|---|---|
| The boy will get a pastry. | The little boy is already up on his toes with one hand on the glass, so he is probably about to ask his mother for one of the cookies in the case. |
| Someone will spill something. | The woman in the green apron is carrying two full cups on a tray through a narrow gap behind the counter, so she will have to wait for the man with the pastry tray to move before she can get past. |
| The queue will get annoyed. | The queue already reaches the door, so once this customer finishes paying the staff will probably start serving two people at a time to clear it. |
The right-hand column is not longer because it is padded. It is longer because it contains the evidence, and the evidence is the part being scored.
#A CELPIP Speaking Task 4 template you can adapt
Two predictions, each with its evidence, then one short closing thought. That fills 60 seconds comfortably without rushing.
Looking at this picture, I think two things are about to happen.
First, [prediction one]. I think that because [what you can see that supports it].
Then [prediction two], since [what you can see that supports it].
After that, [one short consequence or resolution].
So overall, I would expect [a one-line summary of how the moment ends].
The closing line does real work. It gives the answer an ending rather than a sudden stop, and it is the easiest place to show a little range of language without straining for it.
#How many predictions should you make?
Two, developed, beats four listed. About 93% of the responses we looked at made more than one prediction, so this is not where most people lose marks, but it is where the ceiling sits.
Two predictions with evidence gives you roughly four sentences of real content, which is about right for 60 seconds. Four predictions with no evidence gives you a list, and a list has nothing in it for the rater to reward under Content and Coherence.
If you finish early, do not add a third prediction. Extend the second one: say what happens after that, or how the people involved will probably react.
#Two worked answers
Both answers below are about the bakery scene pictured above, so you can check each claim against the image yourself.
I think the shop will get very busy and the staff will not be able to serve everyone. Then the customers will start complaining because they are waiting too long, and some of them will get angry and leave without buying anything. The manager will probably come out and apologise to everybody. Maybe they will have to close the shop early because they run out of bread. It could be a really bad day for the bakery.
Why it loses marks: the first clause is close to a prediction. Everything after it is invention. No manager is visible, nobody looks angry, and the shelves are full. The ending is a small story the picture does not support, and a rater following along has nothing to check any of it against.
Looking at this picture, I think two things are about to happen.
First, the little boy at the front is going to ask for something from the display case. I think that because he is already up on his toes with one hand flat against the glass, and he is looking at the cookies rather than at the person he came in with, so he has clearly decided what he wants.
Then I think the woman in the green apron is going to have to stop and wait. She is carrying a tray with two full coffees on it, and the man with the tray of pastries is coming the other way through the same narrow gap behind the counter, so one of them will have to give way.
After that, the customer in the yellow coat will finish paying, because she is already handing her receipt across, and the queue behind her will finally start moving.
So overall, I would expect the next minute to be busy but perfectly normal.
Why it works: every prediction is followed by a because or a since, and each one points at something a listener could verify in the picture. The final line closes the moment instead of trailing off, and choosing busy but perfectly normal over a dramatic ending shows the reading of the scene is honest.
#What are the 30 seconds of preparation for?
Finding the moment that is mid-motion. Nothing else.
Every Task 4 scene contains one thing that cannot stay as it is: something tilting, spilling, reaching, arriving, or about to close. Find it in the first ten seconds and your first prediction writes itself.
Spend the remaining twenty seconds on the second prediction and the evidence for both. You do not need to plan your closing line: by the time you get there, the obvious ending will be in front of you.
Because Task 4 reuses the Task 3 picture, you also start with an advantage nobody else on the test gets. If you described the scene carefully a minute ago, you already know where everything is. Read the guide on describing a scene and Task 4 gets easier as a side effect.
#How much do you need to say in 60 seconds?
There is no word requirement, since it is a spoken task. As a guide, Task 4 answers on HelloCelpip average about 130 words across the 60 seconds, which is an ordinary, unhurried pace.
That is enough for two predictions with evidence, one consequence and a closing line. If you are running out of time, it is almost always because you described the scene again before predicting anything. The rater has already heard your description. Start at the future.
#What do CELPIP raters listen for?
Every Speaking response is rated in four areas, published in the CELPIP Speaking Performance Standards.
Across more than 12,000 anonymised, aggregated Task 4 responses on HelloCelpip, about a third of everything our evaluator flags is a vocabulary problem, which is the wrong word or a phrase that is not how a native speaker would put it. The other three areas sit close together behind it. This is the useful thing about Task 4. Because most learners already do the task correctly, the way to improve here is not a change of strategy. It is the ordinary work of better word choice and steadier delivery, which pays off across all eight tasks rather than just this one.
#Common Task 4 mistakes
- Describing the scene again. The rater heard your description in Task 3. Predictions only.
- Predicting things nothing in the picture supports. Roughly one answer in eight drifts into invention.
- Listing four predictions with no evidence instead of developing two.
- Using only will. Vary it: is going to, will probably, is likely to, I would expect.
- Forgetting the word probably. The question asks what will most probably happen, so hedging is correct here rather than weak.
- Stopping mid-thought when the timer feels close. A one-line summary is a better ending than a sentence cut in half.
#How to practise Task 4
Practise the two tasks as a pair, because that is how they arrive. Describe the scene for 60 seconds, then predict from the same picture for another 60. It doubles the value of every image and it trains the habit of carrying detail forward.
Then check one thing in the recording: count how many of your predictions were followed by evidence. If the answer is fewer than all of them, that is the whole fix.
#Frequently asked questions
What is CELPIP Speaking Task 4?
How long do you get for CELPIP Speaking Task 4?
Is the Task 4 picture the same as the Task 3 picture?
What do learners get wrong most often in Task 4?
How many predictions should I make in Task 4?
What tense should I use in Speaking Task 4?
How many words should a Task 4 answer be?
Confirm the eight Speaking tasks and the 15-minute Speaking section on the official CELPIP Test Format page. Per-task timing of 30 seconds to prepare and 60 seconds to speak is published in CELPIP’s own Speaking Pro study pack, and the four rated areas come from the CELPIP Speaking Performance Standards. Your HelloCelpip level uses the same 0 to 12 CELPIP scale, so you can track real progress between now and test day. Your official score comes from Paragon Testing Enterprises on test day. HelloCelpip is an independent study resource and is not affiliated with CELPIP or Paragon Testing Enterprises.
Official sources:
Last updated: July 27, 2026