Speaking

CELPIP Speaking Task 3: Describe a Scene Clearly

Sep 13, 2026

In CELPIP Speaking Task 3, help a listener understand a scene by giving an overview and describing selected people, objects, and actions in a clear spatial order. You do not need to mention everything. Your priority is to create an understandable mental picture using accurate details and language that connects those details.

The official Listening and Speaking Study Guide gives this task thirty seconds of preparation and sixty seconds of recording time. Follow the instructions shown in your test. The official test format identifies Task 3 as Describing a Scene. All scene exercises below are original written descriptions for language practice; they are not supplied pictures, official test images, or substitutes for practising visual observation with actual images.

A busy illustration can make speakers feel that every detail is equally important. That pressure often produces a rapid list with no clear organization. A more effective approach is to decide what the overall scene shows, choose a route through it, and describe a manageable set of actions with enough detail for a listener to locate them.

Begin with the scene's overall identity

Your opening should orient the listener. “This appears to be a busy community kitchen where several people are preparing food” establishes a place and a shared activity. It gives later details a framework. Without that overview, descriptions of bowls, aprons, and tables may remain disconnected for several sentences.

Use the level of certainty supported by the picture. If the setting is clearly a playground, say so directly. If it could be a school hall or community centre, use a broader description such as “a large indoor room set up for an event.” You do not need to invent a precise location to sound confident.

Avoid unsupported names, dates, and purposes. A building with books and desks may be a library, but the picture alone may not establish which city it is in or whether an official ceremony is taking place. Describe what the image supports. The task rewards your ability to communicate a scene, not your ability to supply a fictional backstory.

An overview should be brief enough to leave space for the visible activity. Compare “This is a very beautiful, interesting, colourful and wonderful picture of a place where there are many different people” with “The scene shows a crowded outdoor market.” The shorter opening gives the listener more useful information and creates a natural starting point.

If you cannot identify the exact setting, begin with its visible characteristics. “Several people are gathered around long tables in an outdoor area” is still an informative overview. You can clarify the likely setting later if other details support it. Do not spend the entire preparation period searching for a perfect location label.

Choose a route that the listener can follow

Spatial organization turns individual observations into a scene. You might move from foreground to background, left to right, or from the main activity outward. Choose the route that fits the image. No single route is required, but frequent unexplained jumps make the listener work harder to connect the details.

Imagine an original written scene description of a park activity day. Near the front, two children are decorating a sign. In the centre, an adult is demonstrating how to plant seeds. On the right, several people are waiting beside a table of small pots. In the background, volunteers are carrying empty boxes toward a van.

A foreground-to-background route works naturally here because the activities occupy clear layers. You could describe the children, move to the central demonstration, mention the waiting group, and finish with the volunteers. Repeating the word “behind” occasionally is less harmful than moving randomly among these groups without locating them.

Left-to-right organization may work better for a long street scene or a row of shops. A central-focus route suits an illustration dominated by one event, such as a group gathered around a broken display. Start with the main event, then describe nearby reactions and supporting details.

Treat your route as a guide rather than a rigid obligation. If the most important action is in the background, you can begin there after the overview. Make the transition explicit: “The main activity seems to be at the back, where…” What matters is that the listener understands where your attention is moving.

Describe people through actions and identifying details

A useful person description answers three questions: which person, where, and doing what? “A woman is working” answers only the last question, and even that answer is vague. “Near the front counter, a woman in a striped apron is arranging cups on a tray” gives the listener a distinct person and action.

Clothing, position, and an object being handled can identify people without guessing who they are. “The person holding a clipboard” may be more useful than “the manager” when the picture does not establish an occupation. A clipboard suggests a role, but it does not prove a job title.

Use relationships carefully. Two people standing together might be friends, coworkers, or strangers. Unless the prompt or image clearly supports the relationship, describe their visible interaction: “Two adults are talking beside the entrance.” This is more accurate than declaring that a husband is arguing with his wife based only on their positions.

Present continuous is useful for visible actions: “is carrying,” “are waiting,” and “is pointing.” Combine it with precise objects and locations. Instead of “A man is doing something with a thing,” say “A man is tightening a strap around a large box.” If you do not know the object's exact name, describe its appearance or function.

Not every sentence needs to begin with “There is.” That structure is useful for introducing an object, but you can then move to active sentences: “A volunteer is handing out maps near the gate. Several visitors are checking the routes before entering.” The language becomes more connected because the second sentence develops the same area of the scene.

Original worked description: a repair workshop

This is a text-based scene drill. Imagine an illustration of a community repair workshop. A long table runs across the centre. On the left, a person in a green shirt is holding a lamp while another person checks its cable. At the front, a child is placing loose screws into a small container. On the right, an adult is adjusting a bicycle wheel. Two visitors are waiting near the rear doorway, one holding a broken toaster. Shelves of tools line the back wall.

An original practice response could sound like this:

The scene appears to show a community workshop where people are repairing household items. A long worktable fills the middle of the room, and several activities are happening around it.

On the left, a person in a green shirt is holding a lamp steady while someone beside them examines the cable. Near the front of the table, a child is collecting small screws and putting them into a container, perhaps to keep the work area organized.

Over on the right, another adult is leaning toward a bicycle and adjusting its front wheel. In the background, two visitors are waiting by the doorway. One of them is carrying a toaster that may need repairing.

There are shelves of tools along the back wall, so the room looks well equipped for several kinds of repair work.

The response moves through the scene in a predictable route and identifies each group by location. Actions carry most of the information. The lamp is being held and examined, screws are being collected, and a wheel is being adjusted. These verbs help the listener visualize activity rather than merely register that people and objects exist.

The response also separates observation from interpretation. “Perhaps to keep the work area organized” is a modest inference about purpose, and “may need repairing” marks uncertainty about the toaster. Neither statement replaces the visible description. If you removed those interpretations, the scene would remain understandable.

Use spatial language accurately

Spatial phrases should help the listener locate details, not function as decorative vocabulary. “In the foreground” refers to the area that appears closest to the viewer. “In the background” refers to the more distant area. “In the centre” identifies a central position, while “on the left” and “on the right” usually refer to the viewer's perspective.

Be consistent about that perspective. If you suddenly describe a person's own left side without making the change clear, the listener may place an object incorrectly. “To the left of the doorway” is generally easier to follow than “on his left” when several people are present and the reference is uncertain.

Use relative positions to connect new details to established ones. Once you introduce a table, you can say “behind the table,” “beside it,” or “at the far end.” The listener now has an anchor. Creating a chain of clear anchors is more efficient than repeatedly restarting with isolated descriptions.

Be careful with “between,” “among,” and “in front of.” A chair between two desks has a specific position. A person among a crowd is surrounded by a group. Someone in front of a window may partially block it from the viewer. Practise these meanings with objects on your own desk before trying to use them quickly in a complex image.

If a precise spatial term is difficult to retrieve, use a simpler accurate phrase. “Near the back” can work instead of “in the distant background.” A familiar phrase delivered clearly is preferable to a more elaborate phrase that makes the listener uncertain about the location.

Separate visible facts from reasonable impressions

Pictures can suggest emotions and purposes, but they rarely prove them. A person with raised eyebrows and an open mouth may look surprised. A person looking down may be reading, searching, or concentrating. Use cautious language when you move beyond what is directly visible.

Compare “The woman is furious because her employee has ruined the event” with “The woman appears upset and is pointing toward the fallen display.” The second sentence describes expression and action without inventing a relationship or cause. It gives the listener concrete evidence and leaves uncertain details appropriately uncertain.

You do not need to hedge every observation. “There are three chairs beside the table” can be direct when the chairs are clearly visible. Overusing “maybe” and “perhaps” for obvious details can make an otherwise clear response sound hesitant. Reserve uncertainty language for actual interpretation.

Present description and prediction also need different attention. Task 3 focuses on what is happening in the scene. A brief inference may help explain an action, but a long account of what will happen next shifts away from the task. Save developed future events for Task 4 prediction practice.

During review, mark statements as visible, inferred, or invented. Visible statements describe the image. Inferred statements are plausible interpretations supported by cues. Invented statements have no clear support. Remove inventions unless the prompt explicitly requests imaginative development, and make sure inference never overwhelms the actual description.

Build vocabulary around action families

Learning lists of unusual nouns is less useful than building a flexible vocabulary for common actions. Many scenes include people carrying, placing, handing, reaching, pointing, leaning, bending, arranging, repairing, or waiting. These verbs work across parks, stores, classrooms, stations, and community events.

Learn the patterns that accompany the verbs. Someone hands an item to another person, reaches for an object, leans against a wall, or bends down to pick something up. Knowing the preposition and sentence pattern helps you retrieve the whole action accurately. A verb alone may not be enough during a timed response.

Distinguish similar movements. Carrying means transporting an object while supporting it. Holding does not necessarily involve movement. Lifting emphasizes moving something upward. Dragging suggests pulling it along a surface. Choosing among these familiar verbs can make your description substantially more precise.

Adjectives are most useful when they identify or distinguish. “A small red bag” can separate one object from others. “A very nice, beautiful bag” adds evaluation but may not help visualization. Select size, shape, colour, material, and position when those features matter to the listener.

Practise circumlocution for unknown objects. You might say “a folding stand with three legs” instead of an unavailable technical word. Describe what the object looks like, where it is, and how someone is using it. This approach keeps the scene moving and demonstrates usable communication even when your vocabulary has a gap.

Repair a scattered description

Here is a deliberately weak partial response to the workshop drill: “There is a bicycle. There are people. A child is there. There is a lamp. The room is good. A man is waiting, and tools are there.” It mentions several accurate details, but the listener cannot easily connect them or imagine the arrangement.

The first repair is an overview: “Several people are repairing items in a shared workshop.” The second is a route: begin at the central table and move outward. The third is action detail: explain that one person is checking a lamp cable, a child is collecting screws, and another person is adjusting a wheel.

A revised partial description could be: “At the central table, two people are examining a lamp, while a child in front of them collects loose screws. On the right, another adult is adjusting a bicycle wheel.” This version uses fewer separate starts and provides stronger relationships between people, locations, and actions.

Notice that the repair does not depend on rare words. The improvement comes from selection and organization. Many speakers assume their main problem is vocabulary when the listener is actually struggling with disconnected references. Review the structure before deciding that you need a large new word list.

Also check pronouns. “He is helping him while he holds it” becomes difficult when several people and objects are present. Repeat a short identifying noun when needed: “The volunteer is helping the visitor hold the lamp.” A little repetition is preferable to ambiguity that makes the scene impossible to reconstruct.

Use preparation time to select, not catalogue

During the preparation period, first identify the overall setting. Then locate two or three areas with actions you can describe confidently. Choose an order and note a few useful verbs. You do not need a written inventory of every object, and you do not need to plan a sentence for every person.

An efficient note set for the workshop might read “repair room / left lamp cable / front screws / right wheel / back waiting.” Each note has a location or an action. A list such as “green, red, tall, small, nice” would be less useful because it does not preserve the scene's structure.

Select details partly by importance and partly by describability. A central unusual action may be worth the effort to explain. A tiny background object with an unfamiliar name may not deserve much time. The goal is a coherent representation of selected features, not an exhaustive visual report.

If the illustration is crowded, group similar activity. “Several visitors are waiting near the entrance” can summarize a queue without describing every person. Then focus on one distinguishing detail, such as a visitor carrying a large box. Grouping protects time while maintaining the sense of a busy scene.

Plan an ending that completes your route. It might describe the background or summarize the atmosphere through evidence: “Most people are focused on their tasks, and the room seems busy but organized.” Avoid a generic closing such as “That is all I can see” when a final useful observation would fit naturally.

Practise observation separately from speaking

Text scene drills help with vocabulary and organization, but real picture practice is necessary to develop visual selection. Use an image you have permission to view or an official free practice resource. Give yourself a short observation period, hide the image, and list the main actions you remember.

Compare your notes with the image before recording. This untimed check can reveal whether you are noticing actions or only objects. If your list contains table, chair, window, and person but no verbs, deliberately search for interactions: who is passing something, looking toward something, or reacting to someone else?

Next, describe the image while following one spatial route. Record without restarting. On review, draw a rough map based only on your spoken description. If the locations are impossible to place, your spatial references need work. The drawing does not need artistic quality; it simply tests whether the words supply usable relationships.

For a partner exercise, let one person describe an image while the other writes a brief scene summary without seeing it. Then compare the summary with the original. Ask which details were clear and which references were confusing. This gives more useful feedback than asking whether the response sounded impressive.

Finally, repeat the same scene with a different route. Moving left to right after a foreground-to-background attempt forces flexible organization. You are learning to communicate visual relationships rather than memorizing one description. Keep the factual details stable while changing the sequence and wording.

Review clarity using the official dimensions

The official results page lists Content/Coherence, Vocabulary, Listenability, and Task Fulfillment for Speaking. In a scene description, content and coherence involve the selection and organization of visible details. Vocabulary involves accurate names, actions, and relationships. Listenability concerns how easily your delivery can be followed, and task fulfillment includes responding to the actual picture-description request.

A practical review begins with three content checks: the setting is established, the route is traceable, and the main actions are understandable. Then choose one language issue, such as repeated “there is” openings or unclear pronouns. Correcting one recurring issue across a recording can produce a noticeable improvement without overwhelming your practice.

Listen for pauses inside important phrases. “A woman is handing…a small…box…to another person” may be harder to follow than a short pause before the complete action. Practise saying the person-action-object unit together at a comfortable pace. Do not rush merely to include another background detail.

When you finish, state one measurable goal for the next scene: introduce the setting immediately, use three accurate spatial phrases, or replace vague verbs with precise actions. These are practice targets, not scoring formulas. They help you evaluate a specific improvement while keeping the overall purpose of the task in view.

Frequently asked questions

Must I describe every person in the picture?

No. The supplied guide encourages selecting details rather than trying to cover everything. Describe enough meaningful activity to create a coherent scene. Group similar background actions when appropriate, and spend more time on details you can explain clearly and accurately.

Can I guess how people feel?

You can describe a reasonable impression when visual cues support it, using language such as “appears worried” or “seems pleased.” Avoid inventing motives, relationships, or past events. Visible expression and action should remain the foundation of the description.

What if I do not know an object's name?

Use familiar language to explain its appearance, location, or use. A description such as “a long-handled tool for sweeping the floor” can keep your response understandable. Do not stop the whole description to search for a single technical noun.

Should I practise from written scene descriptions?

They are useful for language organization, but they do not reproduce the visual task. Combine text drills with actual image observation and timed speaking. Written descriptions have already selected the details for you, so image practice is needed to develop that selection skill independently.

Sources and further reading