AN INTERACTIVE DIGITAL ART INSTALLATIONINDEPENDENT / BROWSER-BASED

THE RAIN
LISTENS01

A moment between
motion and stillness.

A rain field that responds to a hand.
Open your palm. Let the falling world pause.

SCROLL TO EXPLORE ↓
GENERATIVE GRAPHIC / NOT AN EXPERIENCE CAPTURE
CASE STUDY / 01—08InteractionVisual systemChallengesEvidence
01 / OVERVIEW

A small gesture.
A change in the weather.

THE RAIN LISTENS is a browser-based interactive art prototype. It brings a live camera image into a three-dimensional rain field, letting the viewer change its direction and rhythm through hand gestures.

The aim is to place a person among droplets held in time. A screen and webcam are the entry point. “Installation” describes the artistic format; a physical exhibition deployment has not been documented.

FORMAT
Interactive digital art
APPROACH
Concept, visual iteration & AI-assisted development
PLATFORM
Desktop Chrome / Edge
RELEASE
Cloudflare Pages / HTTPS
01 / HERO DEMOLIVE WEBGL / 00:36
Actual WebGL runtime capture / 6,500 droplets / Cinematic Final experimental preset. Fall → suspend → repel → reverse → resume. No camera or audio; scripted controls. This does not validate real hand tracking or person occlusion.
02 / THE CONCEPT
MOTION / TIME / NATURE

Rain keeps falling.
A person can choose
to stay a moment.

Rain makes time visible through direction, speed and repetition. As the fall slows, weather becomes a field of individual droplets that can be observed.

An open hand needs little explanation. It gives the body a direct relationship with the surrounding field. The roughly three-second deceleration leaves enough time to see the change happen.

03 / THE INTERACTION

The body becomes
the control.

No contact with the screen is required. The system prioritises one hand and uses a confirmed gesture to control the rain.

OPEN PALM — SUSPEND

The rain decelerates over roughly three seconds. With the hand held still, residual drift settles to exact rest.

01

FIST — REVERSE

The field transitions into upward motion. Falling and rising speeds are independently adjustable; switching keeps the current particle positions.

02

HAND MOVEMENT — REPEL

Moving an open palm creates a soft local repulsion field. Its position and strength are smoothed to reduce jitter and abrupt deflection.

03

NO HAND — RESUME

After approximately 700 ms without a hand, natural falling resumes. An ambiguous OTHER classification does not trigger a new action.

04

Without a camera: R / fall, F / suspend, B / reverse; the mouse controls repulsion. Manual control pauses hand tracking until it is resumed through Help.

04 / VISUAL SYSTEM

As speed disappears,
the droplet appears.

Streak length follows actual particle velocity. As speed approaches zero, the streak continuously shortens into an elliptical droplet. Reversal passes through the same transition.

Baseline remains the initial visual. The optional Cinematic Final preset combines continuous perspective with analytic surface normals, local highlights and edge reflections. Its photographic quality still awaits human acceptance.

VELOCITY → FORM
Interactive schematic / not the production shader or a runtime capture. The switch illustrates the relationship between speed and form.
Rain in motion
Rain in motion / Experience screenshot
Rain at rest
Rain at rest / Experience screenshot

Size comes from distance.

A 52° perspective camera projects a continuous depth volume. In Continuous Perspective, equal physical diameters naturally shrink with distance, without fixed screen-size compensation. Near, middle and far are statistical bands.

Equal-diameter geometric samples in the experimental mode: 720 CSS px viewport height, diameter .243 world units. World units are not calibrated metres; these are not visibility statistics for the whole field.

5 units35.87 px
21 units8.54 px
37 units4.85 px
Illustration shown at 2× / data from the projection report.
CAMERA + BACK RAIN × (1 − PERSON ALPHA) + FRONT RAIN

A person alpha mask occludes background rain while foreground rain remains visible over the live, naturally coloured video. A fixed z = −7 plane approximates the person; segmentation supplies no real body depth. Droplets are additive transparent billboards, not true video refraction.

05 / TECHNICAL SYSTEM

One image.
Two inference paths.

Camera capture, vision inference and rain rendering have separate lifecycles. Frames pass through in-browser Worker messages; no backend is involved.

USER ACTION

CameraSession

getUserMedia · video only

DISPLAY PATH

Mirrored video

Shared cover / crop transform

↓

Transparent WebGL rain

Three.js + GPU person-mask sampling

LOCAL INFERENCE / TWO WORKERS

Hand Landmarker

21 landmarks → rules → stable gesture

ImageSegmenter

Confidence alpha → temporal filter → R8 mask

React state + controllers

FALL / SUSPEND / REVERSE → motion integration

Web Audio

Actual motion energy → filtered noise / gain

Architecture derived from current source. Inference input is neither mirrored nor cropped; display, hand position and mask share coordinate transforms.

React / Vite / TypeScript

React hooks hold experience and Studio state. Local controllers such as CameraSession, GestureFilter and DepthFlow manage timing; no Redux or extra state framework is used. Vite bundles the app, Workers and assets.

Three.js / WebGL

Mesh + InstancedBufferGeometry draws 6,500 instanced quads. Motion integration runs on the CPU; shape and occlusion run on the GPU. The implementation does not use the Three.InstancedMesh class. Low mode draws 2,600 while retaining all motion state.

MediaPipe / Web Workers

Both models use the CPU delegate, with one frame in flight per Worker. Hand tracking takes priority; segmentation lowers its rate and resolution under hand or render pressure. Failures preserve keyboard and mouse interaction.

Web Audio / lifecycle

Locally filtered noise supplies rain sound. Gain follows actual motion energy and falls to silence at rest. Sound requires explicit activation. Exit releases tracks, Workers, textures, geometry, materials and AudioContext. Ordinary tab hiding pauses work; camera tracks may remain active.

06 / DESIGN CHALLENGES

Invisible rain needed
more than brighter pixels.

01

Readability before material.

A typical Baseline droplet could project to only .28 CSS px in the distance. Soft contours and compounded opacity weakened coverage further. An early screen-size clamp improved readability but weakened perspective. Later experiments used a shared physical scale and nearer continuous depths, rather than more particles.

02

A two-dimensional person in three-dimensional rain.

Segmentation gives a silhouette, not body depth. Shared mirror and cover transforms solve alignment; fragment masking establishes foreground and background. The fixed person plane remains an approximation. A darkening video filter and overlays were removed, making natural colour the default.

03

A classification is not yet an interaction.

Open palms and fists were initially difficult to trigger. Rules now combine joint angles, palm-normalised distances and extension evidence while allowing natural thumb poses. Confirmation, hysteresis and brief tracking-loss tolerance reduce flicker. A five-second diagnostic separates detection, raw classification and stable state.

04

Stillness has to finish.

Pure exponential decay leaves a tiny residual speed. A finite-time curve settles exactly at around 2.92–3.08 seconds, with only 160 ms of depth staggering. A still hand stops driving the field; residual drift reaches zero. Interrupted transitions continue from the current velocity.

05

A real-time budget, and an end to every resource.

Inference is decoupled from rendering; Workers never queue frames, and segmentation yields budget to hands. AudioContext creation can briefly stall. Some IAB top-level tests showed low rAF frequency, so successful iframe short runs are not a universal performance guarantee. Cleanup also guards against late permission results reopening a camera.

06

Keep the camera experiments. Default to a fixed frame.

Dolly, Truck and Camera Reveal explore space after the freeze. The rain uses true camera translation, but the person remains a flat video image; a centre crop cannot create new viewpoints. The experiments stay in Studio. Audience disables them and restores fixed framing.

07 / PROCESS

From a working sketch
to a release with evidence.

  1. STAGE 1

    Camera + Rain MVP

    Mirrored camera, instanced rain, keyboard states, mouse repulsion and cleanup.

  2. STAGE 2 / 2.5

    Hand tracking + diagnosis

    Local single-hand inference, geometric rules, confirmation, skeleton overlay and five-second diagnosis.

  3. STAGE 3A / 3B

    Person segmentation + occlusion

    Low-resolution alpha, shared mapping, GPU foreground/background compositing and stale-mask fallback.

  4. STAGE 4A—4E.3

    Motion, light + visual studies

    Material and motion comparisons, true perspective, camera brightness correction, then dolly and close-up reveal experiments.

  5. STAGE 5A / 5B

    Cinematic preset + audience experience

    Optional Final preset, minimal Audience, full Studio, permission flow and independent directional speeds.

  6. STAGE 5C.1 / 5D

    Release audit + HTTPS deployment

    Production preview, resource hashes and privacy documentation. Deployed to Cloudflare Pages through Wrangler Direct Upload on 2026-10-09.

The limits of validation belong in the record.

88/88

automated tests passed at release

18 files

online resources verified over HTTPS + SHA-256

6,500

normal-mode instances / 2,600 in low mode

Stage 5B recorded three 60 FPS samples with blank synthetic video, both real models, audio and 6,500 instances. This did not exercise valid hand or person contours and is not sustained human performance evidence. The report provides no hardware benchmark suitable for generalisation.

IMPLEMENTED

Gesture interaction, segmentation occlusion, adaptive sound, Audience/Studio, independent speeds and resource management.

VALIDATED

Automated, GPU and synthetic browser checks; online asset integrity. The user reported human acceptance for Stage 3.5, 4A and 4B; this does not extend to every later or online feature.

EXPERIMENTAL

Continuous Perspective, Cinematic Droplets/Final and Dolly/Truck/Reveal are implemented optional studies, not initial Audience defaults. Photographic quality needs human judgement.

PLANNED

Stage 6A real demo footage, matching comparison assets, and online human and sustained-device acceptance. Not completed.

08 / REFLECTION

Keep visual judgement
in human hands.

This project carries an artistic proposition into an experience that can be used, diagnosed and deployed. Graphics, vision and sound share the same motion state, while each must be able to stop and recover independently.

AI assisted implementation, diagnosis and regression testing. Human feedback redirected droplet readability, spatial structure and camera choices. A render does not establish that a droplet looks convincing; a passing test cannot decide the work’s cinematic quality.

The current choice is to keep natural video and a fixed default frame, while leaving unaccepted visual alternatives available for comparison. The next evidence should be a real person, standing in the rain field and completing the gestures.

Enter the rain ↗

Camera and sound require explicit activation. Keyboard and mouse interaction remains available without a camera.