OPEN PALM — SUSPEND
The rain decelerates over roughly three seconds. With the hand held still, residual drift settles to exact rest.
A moment between
motion and stillness.
A rain field that responds to a hand.
Open your palm. Let the falling world pause.
THE RAIN LISTENS is a browser-based interactive art prototype. It brings a live camera image into a three-dimensional rain field, letting the viewer change its direction and rhythm through hand gestures.
The aim is to place a person among droplets held in time. A screen and webcam are the entry point. “Installation” describes the artistic format; a physical exhibition deployment has not been documented.
Rain makes time visible through direction, speed and repetition. As the fall slows, weather becomes a field of individual droplets that can be observed.
An open hand needs little explanation. It gives the body a direct relationship with the surrounding field. The roughly three-second deceleration leaves enough time to see the change happen.
No contact with the screen is required. The system prioritises one hand and uses a confirmed gesture to control the rain.
The rain decelerates over roughly three seconds. With the hand held still, residual drift settles to exact rest.
The field transitions into upward motion. Falling and rising speeds are independently adjustable; switching keeps the current particle positions.
Moving an open palm creates a soft local repulsion field. Its position and strength are smoothed to reduce jitter and abrupt deflection.
After approximately 700 ms without a hand, natural falling resumes. An ambiguous OTHER classification does not trigger a new action.
Without a camera: R / fall, F / suspend, B / reverse; the mouse controls repulsion. Manual control pauses hand tracking until it is resumed through Help.
Streak length follows actual particle velocity. As speed approaches zero, the streak continuously shortens into an elliptical droplet. Reversal passes through the same transition.
Baseline remains the initial visual. The optional Cinematic Final preset combines continuous perspective with analytic surface normals, local highlights and edge reflections. Its photographic quality still awaits human acceptance.
A 52° perspective camera projects a continuous depth volume. In Continuous Perspective, equal physical diameters naturally shrink with distance, without fixed screen-size compensation. Near, middle and far are statistical bands.
Equal-diameter geometric samples in the experimental mode: 720 CSS px viewport height, diameter .243 world units. World units are not calibrated metres; these are not visibility statistics for the whole field.
A person alpha mask occludes background rain while foreground rain remains visible over the live, naturally coloured video. A fixed z = −7 plane approximates the person; segmentation supplies no real body depth. Droplets are additive transparent billboards, not true video refraction.
Camera capture, vision inference and rain rendering have separate lifecycles. Frames pass through in-browser Worker messages; no backend is involved.
getUserMedia · video only
Shared cover / crop transform
Three.js + GPU person-mask sampling
21 landmarks → rules → stable gesture
Confidence alpha → temporal filter → R8 mask
FALL / SUSPEND / REVERSE → motion integration
Actual motion energy → filtered noise / gain
React hooks hold experience and Studio state. Local controllers such as CameraSession, GestureFilter and DepthFlow manage timing; no Redux or extra state framework is used. Vite bundles the app, Workers and assets.
Mesh + InstancedBufferGeometry draws 6,500 instanced quads. Motion integration runs on the CPU; shape and occlusion run on the GPU. The implementation does not use the Three.InstancedMesh class. Low mode draws 2,600 while retaining all motion state.
Both models use the CPU delegate, with one frame in flight per Worker. Hand tracking takes priority; segmentation lowers its rate and resolution under hand or render pressure. Failures preserve keyboard and mouse interaction.
Locally filtered noise supplies rain sound. Gain follows actual motion energy and falls to silence at rest. Sound requires explicit activation. Exit releases tracks, Workers, textures, geometry, materials and AudioContext. Ordinary tab hiding pauses work; camera tracks may remain active.
A typical Baseline droplet could project to only .28 CSS px in the distance. Soft contours and compounded opacity weakened coverage further. An early screen-size clamp improved readability but weakened perspective. Later experiments used a shared physical scale and nearer continuous depths, rather than more particles.
Segmentation gives a silhouette, not body depth. Shared mirror and cover transforms solve alignment; fragment masking establishes foreground and background. The fixed person plane remains an approximation. A darkening video filter and overlays were removed, making natural colour the default.
Open palms and fists were initially difficult to trigger. Rules now combine joint angles, palm-normalised distances and extension evidence while allowing natural thumb poses. Confirmation, hysteresis and brief tracking-loss tolerance reduce flicker. A five-second diagnostic separates detection, raw classification and stable state.
Pure exponential decay leaves a tiny residual speed. A finite-time curve settles exactly at around 2.92–3.08 seconds, with only 160 ms of depth staggering. A still hand stops driving the field; residual drift reaches zero. Interrupted transitions continue from the current velocity.
Inference is decoupled from rendering; Workers never queue frames, and segmentation yields budget to hands. AudioContext creation can briefly stall. Some IAB top-level tests showed low rAF frequency, so successful iframe short runs are not a universal performance guarantee. Cleanup also guards against late permission results reopening a camera.
Dolly, Truck and Camera Reveal explore space after the freeze. The rain uses true camera translation, but the person remains a flat video image; a centre crop cannot create new viewpoints. The experiments stay in Studio. Audience disables them and restores fixed framing.
Mirrored camera, instanced rain, keyboard states, mouse repulsion and cleanup.
Local single-hand inference, geometric rules, confirmation, skeleton overlay and five-second diagnosis.
Low-resolution alpha, shared mapping, GPU foreground/background compositing and stale-mask fallback.
Material and motion comparisons, true perspective, camera brightness correction, then dolly and close-up reveal experiments.
Optional Final preset, minimal Audience, full Studio, permission flow and independent directional speeds.
Production preview, resource hashes and privacy documentation. Deployed to Cloudflare Pages through Wrangler Direct Upload on 2026-10-09.
automated tests passed at release
online resources verified over HTTPS + SHA-256
normal-mode instances / 2,600 in low mode
Stage 5B recorded three 60 FPS samples with blank synthetic video, both real models, audio and 6,500 instances. This did not exercise valid hand or person contours and is not sustained human performance evidence. The report provides no hardware benchmark suitable for generalisation.
Gesture interaction, segmentation occlusion, adaptive sound, Audience/Studio, independent speeds and resource management.
Automated, GPU and synthetic browser checks; online asset integrity. The user reported human acceptance for Stage 3.5, 4A and 4B; this does not extend to every later or online feature.
Continuous Perspective, Cinematic Droplets/Final and Dolly/Truck/Reveal are implemented optional studies, not initial Audience defaults. Photographic quality needs human judgement.
Stage 6A real demo footage, matching comparison assets, and online human and sustained-device acceptance. Not completed.
This project carries an artistic proposition into an experience that can be used, diagnosed and deployed. Graphics, vision and sound share the same motion state, while each must be able to stop and recover independently.
AI assisted implementation, diagnosis and regression testing. Human feedback redirected droplet readability, spatial structure and camera choices. A render does not establish that a droplet looks convincing; a passing test cannot decide the work’s cinematic quality.
The current choice is to keep natural video and a fixed default frame, while leaving unaccepted visual alternatives available for comparison. The next evidence should be a real person, standing in the rain field and completing the gestures.
Camera and sound require explicit activation. Keyboard and mouse interaction remains available without a camera.