Rule-based Schmitt Trigger vs. TinyML on-device classification for 200Hz wearable finger gestures

Hi everyone,

We are prototyping a wearable spatial input interface using an ESP32-S3 and finger-mounted LSM6DSOX sensors, delivering discrete keystrokes over BLE HID with a strict sub-20ms end-to-end latency budget.

Currently, our pipeline uses a heuristic state machine: dynamic jerk thresholding, baseline gravity tracking via a low-pass filter, and a temporal lockout window to prevent false triggers during finger recoil/return strokes.

As we scale from a single-node PoC to a 5-finger chording matrix, we are evaluating whether to stick to an optimized deterministic state machine or implement an on-device TinyML classifier (such as a lightweight 1D-CNN or quantized decision forest) directly on the ESP32.

Key considerations:

  • How does TinyML inference window latency hold up against strict <15ms processing constraints?
  • Is supervised classification resilient enough to handle parasitic tendon movement across adjacent fingers without extensive per-user training?

Would love to hear from anyone who has deployed gesture models on fast-moving finger IMUs.

Hello @everbrown first of all welcome to the Edge Impulse community!

Thanks for sharing the details of your project!

Let me ask internally to explore the best recommendation that we can share with you!

Thanks

1 Like

Hi @everbrown

The latency question is really one about the tradeoff between size and accuracy of the model; we can always make a model run at <15ms… The question is whether we can structure one that’s accurate enough.

Re: "optimized deterministic state machine or implement an on-device TinyML classifier " I don’t know it’s one or the other, it’s might be a combo? FSMs are super small and easy for control problems with a classifier supporting transistions.

Another thing I’d look at, not necessarily for full implementation but more inspiration, is state space models. e.g. [2110.13985] Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers in particular as well as the followup [2111.00396] Efficiently Modeling Long Sequences with Structured State Spaces

Is “discrete keystrokes” the model output? What exactly does that look like?
If you can share the high level structure of your (X, Y_true) data that would help :smiley:

Mat

1 Like

Hi Mat,

Thanks for the feedback. (Forgive me in advance for any errors. I am not the most technical and my engineer is not always available. Some of my response is definitely with the help of AI.)

You hit the core architectural tension: size vs. accuracy vs. real-time deadline. To answer your questions directly:

  1. Output Definition Y_{true}: The model/device output is strictly low-level HID scancodes / chord primitives over BLE (e.g., KEY_A, KEY_MOD_NAV, or a raw 2-node chord state byte). We deliberately offload dictionary expansion/steno lookups to the host machine (Plover/custom drivers) so the ESP32-S3 only solves physical intent, not language parsing.

  2. High-Level Data Structure (X):

    a.Input Stream: Synchronous 6-DoF telemetry sampled at 200Hz (T=5\text{ms}) across two IMU nodes (thumb + index) over 400kHz I2C.
    b.Window Framing: Sliding window $X \in \mathbb{R}^{W \times 12} (a_x, a_y, a_z, g_x, g_y, g_z \times 2$), with window lengths W \in [10, 20] (50–100ms) for gesture classification.

  3. FSM + Classifier Hybrid: What you suggested aligns with where we’re heading. The pure FSM handles the hard temporal gate (suppressing the biomechanical whiplash/rebound spike within a 220ms lockout window). Where edge ML becomes compelling is at the transition boundaries: distinguishing deliberate multi-node chord holds from parasitic flexor tendon crosstalk (e.g., thumb moving sympathetically during an index snap).Digging into the state-space models and S4 references you linked, continuous-time representation across irregular IMU tick intervals is a very clean angle.

Once our bench telemetry harness is logged, I’d love to run early traces through Edge Impulse to compare a lightweight quantized 1D-CNN against our baseline FSM.

Thanks @everbrown for sharing more details!

Let us know if you can start experimenting and get the first results in terms of size and latency.

Looking forward to helping you!

Sounds good. The data types you have here is well supported in Edge Impulse and I reckon you’ll spend more time on the DSP choice ( Processing blocks - Edge Impulse Documentation ) than the actual model .

There’s going to be likely some custom post processing between the model and host ( you’ll want a soft output from the model into the dictionary expansion step to give more flexibility to do beam style decoding )

Mat

2 Likes

Thanks Mat and Marc. That DSP framing confirms our direction…focusing on raw jerk feature extraction on-chip and letting host-side drivers handle the probabilistic beam search across chords. Once we have the raw CSV traces, we’ll pipe them into Edge Impulse to profile the DSP blocks.

1 Like

Thanks for the update @everbrown

Looking forward to learning more about your results!