# «The courteous city» — finished keyframe, report

**Deliverables**
- `final/A.png`: finished keyframe, 3312 × 2480 (4:3), no text.
- `final/A-plate.png`: the same frame with no effects, for the 3D artist. It is pixel-identical to `A.png` everywhere except the effect pixels (about 3.9 % of the frame, all of them lines, rings or boxes).
- `contact-sheet.jpg`: approved base | clean plate | final keyframe.
- `process/`: every step, the call log `calls.jsonl`, all prompts, masks and the scripts that drew the effects (`process/work/`). `process/objects.json` holds every boxed object with its box corners in image pixels.

No B variant. Nearly all of the budget went into making A clean. A second crossing position would have meant a second full city pass at a quality I could not guarantee, so I left it out.

## 1. Where the crossing went and why

**The problem.** In the base image the zebra floats in the middle of the junction. It lands on no kerb, is square to no road, and the tactile paving on the near corner leads nowhere.

**The decision.** The crossing now runs straight across **the main diagonal road, on its leg to the lower right (south-east) of the junction**. It goes from the **near corner plaza** (the corner closest to camera, bottom right of centre) to **the right-hand sidewalk** by the street tree.
- The stripes are parallel to the road and the crossing is square to both kerbs.
- Each end meets a kerb ramp with a tactile warning strip. On the near corner this is the **same tactile paving that already existed** in the approved picture, so that orphan detail now has a purpose.
- The crossing has 11 stripes. I painted it in code, in the correct perspective, onto the asphalt the model had cleaned. I did this so the edges are straight and the crossing is geometrically square. The model then re-rendered it together with the rest of the frame.

**Why here.**
1. It keeps the same intersection and camera, and it puts the crossing in the most readable part of the frame: the open, sunlit asphalt in the lower right, on the diagonal of the composition.
2. It creates the classic, real-world courtesy moment. **The robotaxi comes out of the left side street, is stopped at its stop line (the white "T" mark at the mouth of the side street), and is about to turn right onto the main road, across the zebra.** It gives way to the parent and child already on the crossing. This is exactly the right-turn yield that drivers in Almaty often skip, and it is why the robotaxi's patience reads as courtesy rather than as a traffic jam.
3. The rover waits on the near corner plaza, which is where the people are walking to. It faces them and stands well back from the ramp, leaving the landing free.

**Who is where.**
- **Robotaxi** (turquoise Sonata, roof lidar rack, checker livery on the rear half; any plate or badge removed). It is in the right-hand lane of the left side street, parallel to the kerb, nose at the stop line, pointing into the junction. About 2 m of clear asphalt separates its front axle from the stop line.
- **Pedestrians H1 (adult) and H2 (child).** They hold hands in the middle of the zebra and walk toward the near corner plaza.
- **Rover** (black six-wheel chassis, tall turquoise lid, front sensor head). It stands on the near corner plaza a couple of metres before the ramp and tactile strip. It points up and to the right, toward the crossing and the approaching people, and is stopped.

## 2. Everything added: list for layout and animation

"Toward the mountains" means away from camera. The main road runs from the lower right of the frame toward the mountains.

| ID | Object | Where | Direction / animation |
|---|---|---|---|
| C1 | Silver sedan | Main road, lane next to the right-hand sidewalk, mid-distance | Driving **away** toward the mountains; keeps moving and leaves at the far end |
| C2 | Silver sedan | Main road, far section, on the other side of the centre line | Driving **toward** the junction. In the loop it slows and stops at a stop line on the far side of the crossing (add this line in 3D, mirroring the one in front of the robotaxi) while H1 and H2 cross. It is the second car waiting for the people |
| C3 | Silver hatchback | Left side street, the far lane (beside the left corner plaza) | Driving **away** from the junction toward the left frame edge; leaves the frame |
| C4 | Small silver sedan | Far cross street at the foot of the view | Driving **left**, background traffic |
| C5 | Small silver sedan | Far cross street | Driving **right**, background traffic |
| B1 | Cyclist | Main road, on the right-hand kerb edge, beyond C1 | Riding **toward the mountains**, same direction as C1 |
| P1 | Pedestrian | Far-left sidewalk (the far corner of the left block) | Walking toward the junction |
| P2 | Pedestrian | Left corner plaza | Walking away, toward the mountains |
| P3, P4 | Two pedestrians together | Right-hand sidewalk, near the planters | Walking away, toward the mountains |
| P5 | Pedestrian with backpack | Right-hand sidewalk, by the second bench | Walking toward camera, toward the crossing landing (he is the next person who will use the crossing) |
| H1, H2 | Adult and child | On the zebra | Walking toward the near corner plaza; the rover waits for them |
| S1–S3 | Chrome spheres: large (~1.2 m), medium (~0.8 m), small (~0.5 m) | Near plaza, lower left, in front of the planter | Static; they reflect the street |
| S4–S5 | Chrome spheres: ~1.0 m and ~0.5 m | Left corner plaza, against the planters | Static |
| S6 | Chrome sphere (~0.9 m) | Right-hand sidewalk, beside the planter behind the bench | Static |

All the spheres stand beside the walking routes, not in them.

- Every car, person and the cyclist has a white wireframe box: all 12 edges, no labels. The robotaxi and the rover have none.
- Box corners in image pixels are in `process/objects.json`. Use them to match the 3D boxes to the frame.
- **Background:** Kok-Tobe is now a lower, rounded, wooded foothill in front of the snowy Zailiysky Alatau, with the slender TV tower on its summit. It is grey, in the same haze as the range.
- The sky above the mountains is untouched and empty, ready for the logo and caption.

## 3. Path lines and lidar waves in the loop

- **Robotaxi path line.** A flat, translucent turquoise ribbon about the car's width (≈1.9 m). It starts under the front of the car and runs forward in its lane, brightest at the car and fading forward with a soft glow. It stops before the stop line.
  - While the people cross, the ribbon "breathes" (a gentle brightness pulse, about a 2 s period) but does not grow.
  - When H1 and H2 reach the kerb, the ribbon extends forward and curves right onto the main road across the now-empty zebra. The car follows it, and the loop resets.
  - **Nothing ever crosses the zebra while people are on it.**
- **Rover path line.** A narrower ribbon (≈0.55 m). It runs from under the rover toward the ramp and ends about 1 m before the tactile strip and the arriving people.
  - When the people have passed, it retracts and re-extends toward the ramp. The rover crosses after them, or turns along the plaza.
- **Lidar waves.** These are drawn in the image as three thin rings for each hero:
  - Robotaxi: rings centred on the roof lidar, sitting between roof height and bumper height and tilted slightly down as they spread.
  - Rover: smaller rings from its sensor head, near knee height.
  - They are nearly parallel to the ground and fade with distance, and they are cut where they pass behind the car's body. The rover's rings stay clear of the people.
  - Loop: each ring is born at the sensor, expands and fades out. A new ring is born every ~0.6 s, so three rings are always visible and the pulse loops seamlessly at any loop length that is a multiple of 0.6 s.
- The colour is #0DC3D0 everywhere; boxes are white/very light grey.

## 4. Model and settings that worked

- **The model: `sunburst`.** On the first test (a full re-layout at 1440 → 2048), both models re-composed the camera. `flare` moved more, so I used `sunburst` for every other call.
- **Feed the model 2048 × 1536 input.** Feeding it the same size as its output stopped the reframing: SIFT alignment showed less than 1 px of drift.
- **Remove first, then add.**
  1. Remove the old zebra, the people, the old rover and the old arcs (masked, high quality).
  2. Paint the zebra and the stop line in code.
  3. Move the car and add the people and the rover, guided by a layout sketch (masked).
  4. Add traffic, people and spheres, guided by a marked layout image (masked, protecting the heroes and the sky).
  5. Kok-Tobe plus a fix to C3's direction (masked, with the Kok-Tobe reference).
  6. Rebuild the rover (masked, with the Yandex rover photo). That call greyed out the robotaxi, so I composited in only the rover.
- **One final pass at `max` quality (3312 × 2480),** with the approved original as a material reference. The output was re-aligned to the working frame by homography (the model had zoomed about 2 %).
- **Effects drawn in code (Python/OpenCV),** as the brief allows: path ribbons, lidar rings and wireframe boxes.
  - The ground plane and camera were solved from the painted crossing (about 69° field of view, camera about 6 m high), and the robotaxi's wheelbase confirmed the scale.
  - The people's boxes are fitted in that camera.
  - The car boxes are built from each car's own wheel line and bumper line. This keeps the edges parallel and snug where the generated image is not one strict perspective.
- **Calls used: 10 of 36.** After call 10, the fal status endpoint started returning HTTP 403 on every request (submissions still went through, results could not be fetched). No further model calls were possible, so the last fixes (plate removal, compositing) were done by hand in code.

## 5. Still imperfect: fix in 3D

1. **Scale of the grey cars.** Measured against the road and the robotaxi, the silver cars (especially C2, C4 and C5) are drawn about 0.5–0.65 of true size. Build them at real scale in 3D; the positions and directions in the table are what matter.
2. **Perspective is not strict everywhere.** The generator's image does not follow one exact camera, particularly on the left side street, where C3 is rotated a little toward the camera compared with the kerb. In 3D, align C3 to its lane.
3. **Stop lines.** Only the robotaxi's stop line is painted. Add one for C2's direction on the far side of the crossing, and give lane arrows and road markings a general clean-up in 3D.
4. **The robotaxi's rear bumper.** A licence plate had to be removed by hand, so a small patch on the rear bumper is flat dark grey. The badges and plates on the grey cars were removed or flattened the same way. Build them with no plates or badges.
5. **The rover.** The lid is a slightly boxy interpretation of the Yandex rover, with a small sensor head at the front. Model it from `context/vehicles/483.jpg`. Its lid is a little lower than the real one.
6. **Pedestrians.** The chrome figures have clean anatomy at 100 %, but the small far figures (P1, P3, P4) are soft. The child H2 and the adult H1 hold hands correctly.
7. **Kok-Tobe hill.** The hill silhouette is plausible but generic. Use the real Kok-Tobe profile (`context/refs/koktobe/`).
8. **Light.** The overall tone is marginally darker and more contrasty than the approved base, from the final max render. It stays within the monochrome look, but match the grade to the original in 3D and compositing.
9. **Ribbon length.** The robotaxi's path line is short (the car is stopped close to its line). If the art director wants a longer ribbon in the still, move the car back about 3 m in 3D rather than letting the line reach the zebra.

## Revision 1 (A2)

**Files:** `final/A2.png` and `final/A2-plate.png` (3312 × 2480). They are pixel-identical except for the effect pixels (about 4.3 % of the frame: boxes, ribbons and rings). The updated boxes are in `objects-A2.json`. `A.png` and `A-plate.png` were not touched. Work files are in `process/rev1/`.

**Second rover (R2).**
- It stands on the left corner plaza, on the pale paving between the lamp post (with the two chrome spheres behind it) and the kerb, level with the woman walking there (P2).
- Same robot and materials as R1: black six-wheel chassis, tall turquoise lid, front sensor column, no lettering.
- It is scaled to its distance (about 1.1 m tall, a little below the shoulders of P2 beside it) and has contact and cast shadows matching the bollards.
- It faces up-right, toward the corner and the side-street kerb, and is **stopped**: P2 is walking past just ahead of it.
- Its narrow turquoise ribbon runs on the pavement in front of it and ends before P2's walking line. Three small lidar rings pulse around its sensor head and stop short of her. It has no box.
- **In the loop:** the rings pulse on the same ~0.6 s rhythm as R1 throughout. R2 waits while P2 walks past toward the mountains. When she has cleared its path, the ribbon extends and R2 rolls forward toward the corner, then turns up the left sidewalk, following well behind her. In the 3D layout, R2 should never come within about 1.5 m of a pedestrian.

**Car check and fixes (identity kept: every car and the cyclist was moved or rescaled from its own pixels, none replaced).**
- **C2 was in the wrong lane.** It was driving toward the junction in the right-hand half of the road, which is the north-west-bound side. I moved it one lane toward the left of the frame, into the south-east-bound half, as right-hand traffic requires. I scaled it +12 %, and a model pass repainted clean asphalt and lane dashes where it stood.
- **Scale:** C4 was scaled +45 % and C5 +40 % (both were toy-sized), C3 +15 % and C1 +10 %, each scaled around its own wheel contact so the wheels stay on the ground. The far cars (C4, C5) still read slightly small; build every car at real scale in 3D.
- **Cyclist B1** was riding on the sidewalk paving. I removed him and restored the paving (one model pass on a crop). He is now on the road at the right-hand edge of the north-west-bound kerb lane, just inside the solid edge line, behind C1, riding toward the mountains. He is scaled for his closer position, with his shadow on the asphalt.
- I checked every car: correct direction for its lane, wheels on the ground with contact shadows, no overlaps with kerbs, bollards, other cars or people. All boxes were refitted, and C2, C4, C5 and B1 got new boxes.

**Tooling note.**
- The switched service refused `--quality max` with `not_enough_boost_credits`. So the three new edits (road clean-up, cyclist removal, R2) were rendered at `high` on tight crops. That gives about 1.7× the pixel density of the final frame, so the result matches the `max`-rendered A plate at 100 % zoom.
- Everything else in A2 comes straight from the A plate. Calls used: 13 of 36.

## Revision 2 (A3)

**Files:** `final/A3.png` and `final/A3-plate.png` (3312 × 2480). They are identical except for the effect pixels (boxes, ribbons and rings). A, A2 and their plates are untouched. `contact-sheet-A3.jpg` shows A2 | A3 plate | A3. Work files are in `process/rev2/`.

**What changed.** Every perceived object is now the generic base model of its kind from `context/models-minimal/`, in the Zoox-style material: smooth, softly frosted, milky pale grey-white volumes with no fine detail.
- **C1–C5:** the same generic compact hatchback (CAR sheet).
- **P1–P5 and H1:** the same generic adult (PERSON sheet). P5 lost his backpack.
- **H2:** the generic child (PERSON-CHILD sheet), still holding H1's hand.
- **B1:** the standard cyclist (CYCLIST sheet).

Each object kept its position, wheel or foot contact, direction of travel, pose (walking / riding), scale and shadow.

**What did not change:** the robotaxi, both rovers, their turquoise ribbons and lidar rings, the city, the crossing, the spheres, the trees, the mountains and Kok-Tobe, and the light. Only the pixels inside each object's box area were replaced; everything else is the A2 plate unchanged.

**Boxes:** unchanged. The new bodies sit inside the A2 boxes, so `objects-A2.json` is still valid for A3 and no `objects-A3.json` was needed.

**How it was done.**
- Four `high`-quality edits on crops of the A2 plate (left corner, far road, right sidewalk, crossing), with the model sheets as references. Each result was aligned back by homography and composited only inside the objects' box areas.
- The model drew C4 (far cross street) pointing right, which is the wrong direction. I mirrored its pure side view in code so it drives left again, and a fifth small crop edit repaired the wall and trees behind it.
- Calls used: 18 of 36.

**Known imperfections.**
- Behind C4, the far parapet wall has a ~3 px height step at its right end.
- H1's leading foot touches the edge of his box.
- Both are for 3D to resolve.

## Revision 3 (effects)

**Files:** `fx/A`, `fx/B`, `fx/C`, each with `states.png` (robotaxi + rover 1 close-up at t = 0, 0.5, 1.0, 1.5 s), `state-rover2.png`, `loop.mp4` (full frame, 1280 × 960, 30 fps, H.264, 4 s = two pulses) and `spec.png` / `spec.svg`. `fx/overview.jpg` shows the t = 0.5 s close-ups of A, B and C side by side. `fx/<V>/_t05_closeup.png` is the full-resolution version of that close-up. Everything is drawn by code on `final/A3-plate.png` (no image generation, 0 calls). The scripts are in `process/rev3/`, and `variants.py` holds the single parameter table the spec sheets are printed from.

**How it is built.**
- Every ground effect is evaluated per pixel in the ground plane of the solved camera, so a ring of radius 10 m in the image is a 10 m ring in 3D. The camera was solved from the square zebra and checked against the robotaxi's wheelbase.
- Effects lie on the ground and pass behind everything standing on it (people, the ghost cars and cyclist, bollards, spheres, planters and the heroes themselves). Nothing is ever drawn over a body.
- **Pulse:** 2 s. Robotaxi at 0 s, rover 1 at +0.66 s, rover 2 at +1.33 s.
- **Box flash:** when the wave or sweep reaches a perceived object's footprint, its white box flashes. Attack is 0.06 s, decay 0.30 s; line width ×1.9 with a soft white glow, then it returns to normal. Each hero's hit times are listed on its spec sheet.

**A — Waymo pulse.**
- **Head flash and scan skin:** the lidar head flashes at 0.06 s, and a turquoise dotted scan skin sweeps over each hero's body from front to back (0.03–0.5 s).
- **Ring:** a halftone ring is born tight around the vehicle footprint and expands (robotaxi to 30 m, rovers to 9 m, over 1.6 s, easing out). It thins from 1.6 to 0.35 m and fades out, then rests 0.4 s.
- **Halftone dots:** they sit on a ground grid aligned to the vehicle, pitch 0.22 m (rovers 0.09 m), dot size ∝ √intensity.
- **Ribbon:** it carries a bright 0.9 m band that travels from the vehicle to the ribbon's end in 0.9 s.
- Most "Waymo" and the richest. It is the strongest in daylight because the dots are large near the camera.

**B — Ouster rings.**
- **Rings:** 13 fixed rings around the robotaxi, 8 around each rover, at the radii real lidar beams hit the ground: r = sensor height / tan(beam angle). This gives 3.4 → 31.9 m for the car (h = 1.95 m) and 1.8 → 10.5 m for the rovers (h = 1.1 m).
- **Look:** thin lines made of point returns every 0.6°, with true lidar shadows behind every perceived object.
- **Wave:** a brightness wave runs outward through the rings in 1.2 s.
- **Ribbon:** fine parallel light lines (7 for the car, 3 for the rovers) with a pulse flowing along them in 1.0 s.
- Most technical and calmest. After a first pass looked too faint at full-frame size, the ring base opacity was raised to 0.42 and the lines widened to 0.08 m (rovers 0.045 m). It is still the most delicate of the three on a sunlit LED screen.

**C — Zoox disc.**
- **Disc:** a translucent disc (robotaxi 11 m, rovers 4 m; fill opacity 0.055, bright 0.85 rim) around each hero footprint.
- **Sweep:** a bright sweep edge with a trailing wedge rotates one turn per pulse (180°/s), starting from the heading.
- **Boxes:** each box flashes as the sweep passes over it.
- **Ribbon:** solid, with a soft glow that breathes between 75 % and 100 % over the 2 s.
- Clearest reading of "the car sees in every direction". The taxi disc tints a large area of the junction, so keep the fill low.

**Recommendation:** A for the event screen. It reads first in daylight and matches the Waymo idea the art director referenced. The B-style lidar shadows could be borrowed as a detail.

**Known limits for 3D.**
- The ground is treated as one plane, so the 15 cm kerbs are ignored. In 3D, project the effects onto the real kerb and sidewalk heights.
- The occlusion mattes are hand/grab-cut from the plate and slightly generous around the ghost people.
- The head glows and the body scan skin are image-space approximations. Build them on the real lidar mesh and body in 3D.

## Revision 4 (route line)

**Files (all in `line/`):**
- `L1-wait.png`, `L1-go.png`, `L2-wait.png`, `L2-go.png`: full frame, 3312 × 2480, effect A rings, t = 0.5 s.
- `states.png`: the L1 comet at t = 0, 0.5, 1.0 and 1.5 s, waiting state.
- `loop-A.mp4`, `loop-B.mp4`: L1 with effect A rings, and with effect B rings. Full frame, 4 s, 1280 × 960, 30 fps, H.264, waiting state.
- `routes-plan.png` / `.svg`, `spec.png` / `.svg`.

Everything is drawn in code on `final/A3-plate.png` (no image calls). The scripts are in `process/rev4/`: `routes.py` holds the routes and `line.py` holds the line parameters the spec is printed from.

**The line.** The lane-wide ribbons are gone. Each hero now has a thin neon line lying on the ground in the solved camera:
- **Profile:** near-white hot core (38 % of the width), #0DC3D0 body, and a soft turquoise gaussian glow (σ = 2.4 × width).
- **Real widths:**
  - Robotaxi L1: 0.18 m, on the car centre line.
  - Robotaxi L2: two 0.11 m lines at ±0.80 m, the wheel track.
  - Rovers: 0.09 m.
- **Distance:** on screen a line is never thinner than 1.2 px; further away its opacity falls with its true width, so it thins naturally toward the horizon.
- **Shape:** smooth cubic curves with no kinks. Each line starts under its hero and fades out over its last 8 m at the frame edge or the end of the route.
- **Occlusion:** like the rings, it passes behind every person, car, bollard, sphere and hero.

**Routes (right-hand traffic). The 3D artist animates from `routes-plan`.**
- **Robotaxi.** From its stop line in the side street it turns right in a smooth ~6 m radius into the kerb lane of the main road (2.0 m from the kerb). It crosses the zebra and leaves the frame to the south-east (12 m visible).
  - Yielding segment: the part over the zebra, s = 5.6–9.3 m.
- **Rover 1.** Rolls to the near kerb ramp, crosses on the zebra centre line after H1 and H2 have reached the plaza, and goes up the far ramp. It then turns left onto the right-hand sidewalk, 16 m out from the near kerb (between the bollards and P3/P4/P5), and runs toward the mountains to the horizon.
  - Yielding segment: from the near ramp to the far ramp, s = 1.6–17.0 m.
- **Rover 2.** Waits while P2 walks past, then follows her up the left sidewalk to the end of the block at the far cross street, passing P1 about 1 m to his right (24 m visible).
  - Yielding segment: the first 2.5 m, which crosses P2's path.

**States.**
- **Waiting** (the keyframe moment): the yielding segment is at 30 % intensity and 55 % width, and the comet only glimmers through it (45 %). The rest of the route is fully lit.
- **Go:** when the last person has left the path, the segment brightens and widens to full over 1.0 s (smoothstep).
- In the `-go` stills the people are still where the plate has them, for reference. In the animation they are already on the kerb by then.

**Pulse.**
- Period and offsets as before: 2 s, robotaxi 0 s, rover 1 +0.66 s, rover 2 +1.33 s.
- Once per pulse a short bright comet (robotaxi 5 m, rovers 2 m long) runs from under the hero forward in 1.4 s, then rests 0.6 s.
- How far it runs: the robotaxi's comet goes to the frame edge (8.6 m/s). The rovers' comets cover their first 25 m (≈ 18 m/s) and fade over the last quarter of that run.
- Boxes still flash when effect A or B reaches them.

**L1 vs L2.**
- L1 is cleaner and reads as one "Yandex line", matching the press wall and badge. It is my recommendation, and the one used in the loops.
- L2 matches the approved concept's double line and reads more as "vehicle tracks". Near the camera it is richer, but at screen size its two lines merge in the distance.

**Known limits.**
- The ground is one plane: kerb heights (≈15 cm) are ignored, and the plate is not a perfect plane far away. So in the plan the far kerbs are approximate, while the routes, crossing, stop line and people are exact in the camera's ground frame.
- In 3D, lift the lines onto the real kerb ramps and sidewalk heights.

## Revision 5 (thicker route lines)

**Two thickness options (in `line-thick/`).** Everything else is unchanged from Revision 4: routes, yielding states, comet, phases and effect A rings. It is L1 only, one line per vehicle.
- **T2 (×2):** robotaxi 0.36 m, rovers 0.18 m.
- **T3 (×3):** robotaxi 0.54 m, rovers 0.27 m.

Core (38 % of the width) and glow (σ = 2.4 × width) scale with the width, so each line keeps its character (near-white core, turquoise body, soft glow) and is only bolder. The files are:
- `T2-wait.png`, `T2-go.png`, `T3-wait.png`, `T3-go.png`: full frame, 3312 × 2480, effect A rings, t = 0.5 s.
- `T2-loop-A.mp4`, `T3-loop-A.mp4`: 4 s, 1280 × 960, 30 fps, H.264.
- `spec-thick.png` / `.svg`: the updated widths.

**Recommendation: T2.** It reads clearly on the LED screen from the audience distance and still looks like a line. T3 is the boldest and closest to the approved concept's lines, but near the camera it starts to read as a painted road marking rather than a light. The script is `process/rev4/produce_thick.py` (code only, no image calls).
