What's new in microblx (dev)?

January 14, 2026 — updated 2026-09-23

Although quiet on the master branch, a lot is going on on -dev. Here's a summary of what's coming (and will become the long awaited v1.0).

new blocks

lsdb-intf and ubx-dbus

A D-Bus interface block that exposes the full node API (create/remove blocks, connect ports, set/get configs, load USC models, ...) over D-Bus. The companion ubx-dbus CLI provides interactive access from the shell.

The plugin extension lets you register custom D-Bus objects at runtime by loading Lua plugin files, so application-specific interfaces can sit alongside the standard org.ubx.node API.

The block is based on the lsdbus bindings. Essentially, this permits creating coordination interfaces that can cleanly and safely interact with the hard real-time processing without impacting real-time performance.

lfrb

A new lock-free ring buffer iblock based on a minimal internal liblfq implementation which is based on a Vyukov MPMC queue. This replaces the external liblfds dependency which didn't support arm64. Strongly typed, hard-real-time safe.

gpio

Linux GPIO block using libgpiod v2. Ports are created dynamically from the gpios config array — each input line becomes an out-port, each output line becomes an in-port (type unsigned int). Supports active-low polarity, pull bias, and emit-on-change mode.

iio

Linux Industrial I/O blocks via libiio for ADCs, DACs, IMUs, and environmental sensors. Two variants:

  • ubx/iio — polled, synchronous sysfs read each step; also supports DAC output
  • ubx/iio_buf — buffered, hardware samples at sampling_frequency; ptrig drains the kernel buffer

Ports are created dynamically from the channels config; type is always double (physical scaled value).

gps

GPS block that reads from gpsd via its shared memory interface. Emits a single gps out-port of type ubx_gps_data with position, velocity, fix quality, and uncertainty fields.

webgraph

A luablock that serves an interactive browser graph of the running node using React Flow + ELK.js. Blocks are shown as nodes (colour-coded by state) with ports, configs, and iblock edges. The page auto-refreshes and re-runs the ELK layered layout. Requires luasocket and json.lua; JS libs are fetched from CDN on first load. Self-trigger it at a low rate alongside your composition — no separate process needed.

ubx-launch gained a matching -webgraph [PORT] flag to attach one automatically:

$ ubx-launch -c threshold.usc -webgraph

webgraph showing the threshold example composition

This replaces the old webif block, which was removed — it had grown too complex and unmaintainable.

math_double / math_float

Applies any single-argument math.h function element-wise to an input array, with optional per-element mul and add (y = f(x) * mul + add). The function is selected via the func config string ("sin", "sqrt", "log", etc.).

netsink

A luablock-based streaming sink for live-plotting and telemetry (primarily PlotJuggler). It drains input ports and emits one message per sample over the network. Two transports — UDP datagrams or a ZeroMQ PUB socket — and two encodings — JSON or MessagePack — are selectable, all via lua_str globals. Ports are created dynamically from the ports declaration:

{ name="sink", type="luablock:netsink" },
-- ...
{ name="sink", config = {
     lua_str = [[
        ports     = { x="double", v="double" }
        transport = "zmq"          -- "udp" (default) or "zmq"
        format    = "msgpack"      -- "json" (default) or "msgpack"
        uri       = "tcp://*:9870"
     ]],
} },

Ports default to scalars; a "[N]" suffix declares an array port (pos="double[3]"), which serializes as a nested array and is expanded by PlotJuggler into pos/0, pos/1, ...

Both transports go straight through the LuaJIT FFI — no luasocket or lzmq dependency. For RT decoupling, run the sink on its own thread at a lower rate (thread=1, period=100) with deep connection buffers (buffer_len=N); each step drains the full buffer so no samples are lost on the RT side.

$ ubx-launch -c netsink.usc

netsink streaming sin/cos/tan to PlotJuggler over UDP

(Previously called udpsink; renamed now that it does more than UDP — the migration is a one-word type change.)

signal processing: stats, movavg, ewma, mux/demux

A set of small, generic building blocks for composing signal chains. All are runtime-typed: element type and vector length are chosen via the type and data_len configs instead of at compile time, so one block serves all numeric types.

  • ubx/stats — running min, max, mean, std and cnt of the input signal, computed online with Welford's algorithm and emitted as a struct ubx_stat. Vectors are accumulated per element; the output rate can be throttled with stats_output_rate, and skip_first discards N startup samples, so a single transient outlier doesn't ruin min/max and dominate mean and stddev for the rest of the run.
  • ubx/movavg — fixed-window moving average, with a mode config selecting the window aggregate (mean, median, min or max). The median is RT-safe (insertion sort into a preallocated scratch buffer).
  • ubx/ewma — exponentially weighted moving average (y += alpha * (x - y)), the cheap, state-free-ish alternative to a window filter.
  • ubx/mux, ubx/demux — composition glue: mux concatenates nin input ports into one vector, demux partitions a vector onto nout output ports. The optional in_len/out_len configs assign per-port sub-vector lengths, which makes demux double as a slicing block — connect only the outputs you need.

vstore and latch

Two new connection iblocks for cases where lfrb's lock-free queue is paid for but not used. Neither is a default — pick them per connection, once the respective precondition has been checked.

ubx/vstore is a single-slot value store for a connection whose writer and reader are stepped by the same trigger in the same thread. That pair needs neither a queue nor synchronisation: the write always completes before the read begins. Per connection and step lfrb performs four atomic read-modify-writes — not uniformly cheap, e.g. on an ARMv8.0 target without LSE each becomes an outline-atomics helper call — and spreads freeq/usedq/slot/elem over several cache lines, where vstore is one allocation with the payload next to its flag. It keeps lfrb's read contract exactly (0 when there is no unread value) and counts violations of its own precondition on an overwrites port.

ubx/latch is a single-writer, multi-reader latest-value store: a read does not consume, so any number of readers share one latch and the value stays available until the writer replaces it. Concurrency is a seqlock — the writer is wait-free, a reader catching a write in progress retries and then reports no-data rather than spinning — so unlike vstore it is safe across threads. Use it for published state sampled by an unrelated cycle, or for signals only written on change:

{ src="ptrig.overrun_cnt", tgt="overruns.in", type="ubx/latch" },

The catch is the same for both: the return value means a value exists, not a new value arrived. Anything that must act only on fresh input stays on lfrb.

other changes

blockdiagram: direct luablock loading

.usc compositions can now instantiate Lua blocks directly using the luablock:NAME type prefix — no need to manually import the luablock module, create an instance, and configure lua_file separately:

blocks = {
    { name="ctrl", type="luablock:mycontroller" },
},

The name is resolved against the standard microblx block prefixes.

luablock: self-triggering

luablock now supports built-in self-triggering via the thread and period configs. For non-RT management or interfacing tasks this eliminates the need for a dedicated ptrig:

config = { thread=1, period=100 }  -- self-trigger at 100 ms

preinit / preexit life-cycle hooks

Blocks gained two optional life-cycle hooks, preinit and preexit, running before the regular init/cleanup. They add a second structural-mutation point for the rare case where a block must build its own interface from its configuration — e.g. create ports whose number or type depends on a config, or a config whose value drives the creation of other configs. This is what lets netsink above turn its ports declaration into real ports before the connections are wired up.

Most blocks never need this — init remains the place for ordinary setup. Reach for preinit only when interface creation has to depend on configuration. It's available to C blocks and to luablock (just define Lua preinit/preexit functions); the rationale is written up in docs/dev/004.

ptrig: SCHED_DEADLINE support

ptrig now supports Linux EDF scheduling via sched_policy = "SCHED_DEADLINE". Timing parameters are passed via the sched_deadline config (runtime_ns, deadline_ns, period_ns); deadline_ns and period_ns default to the period config when 0. After each chain trigger sched_yield signals budget exhaustion to the kernel. Budget overruns are caught via SIGXCPU and counted on the deadline_throt_cnt out-port. The sched_deadline in-port allows runtime parameter updates without reconfiguration.

ptrig: dynamic period port

ptrig gained a period in-port for runtime adjustment of the trigger period without reconfiguration.

ptrig: trigger latency measurement

How late does a trigger actually fire? overrun_cnt only counts whole periods lost, so a trigger that is consistently 200 us late on a 1 ms period reports zero overruns, and the tstats measure how long the chain ran, not when it started. The quantity the sleep modes exist to reduce was the one without a metric — answering it meant reaching for the kernel's timerlat tracer.

latency_stats = 1 now writes the per-cycle lateness to a new latency_ns out-port and logs min/max/avg at stop, next to the overrun report. Connect latency_ns to a ubx/stats block for the stddev. It is opt-in because it costs one extra clock read per cycle; the first tstats_skip_first samples are excluded from the logged summary, since the deadline grid is only anchored on the first cycle and that outlier dominated the result.

ptrig: hybrid sleep mode and timer slack

sleep_mode gained a third option: hybrid (2) sleeps until busy_slack_ns before the deadline and busy-waits the rest, which absorbs the OS wakeup latency at busy-wait accuracy for a busy_slack_ns / period duty cycle (default 50 us). busy_slack_ns must exceed the platform's worst-case wakeup latency, otherwise the busy phase never runs and the mode silently degrades to sleep_mode=0.

The new timerslack_ns config sets the trigger thread's slack via prctl(PR_SET_TIMERSLACK). Linux defaults non-realtime threads to 50 us of timer slack, which shows up one-to-one as trigger lateness under SCHED_OTHER; realtime policies ignore it.

Also fixed: after a missed deadline ptrig discarded the realigned grid point and fired one trigger immediately at an arbitrary phase. It now waits for the realigned deadline, so an overrun drops the affected tick(s) and leaves the phase of all subsequent triggers unchanged.

cheaper clock reads: ubx_gettime_ns

ubx_gettime returns a sec/nsec pair, which on the CNTVCT time source costs three 64 bit divisions to construct — on a Cortex-A53 the divisions dominate the call, while the counter read itself is a fraction of it. Most callers don't want the pair: the busy-wait loops built one per spin iteration only to compare it, and ptrig converted it straight back.

The new ubx_gettime_ns returns nanoseconds directly; where the counter frequency divides 1e9 evenly (the common 25/50/100/200 MHz parts) the conversion is a single multiply, and the odd 19.2/24 MHz parts keep the division path. Measured on an AM62x at a 500 us period with sleep_mode=2, mean trigger lateness dropped from 1553 to 1506 ns: the busy-wait exits closer to the deadline because each spin iteration is cheaper. This helps the counter-based time sources only — on the clock_gettime fallback the vDSO call dominates and the arithmetic around it is noise. Which is worth knowing, because both hardware time sources still default to off: the out-of-the-box build pays 65 ns per timestamp against 15 ns for CNTVCT, proportional to how much timing instrumentation is enabled.

const: runtime in port

ubx/cconst and ubx/iconst gained an in port. Writing to it updates the held value at runtime — without a reconfigure cycle. Example use-case: update dynamic configuration safely via the lsdb-intf block.

runtime-typed saturation, ramp and rand

The compile-time typed block variants are being replaced by single, runtime-typed blocks that take the element type and vector length from the type and data_len configs:

  • ubx/saturation replaces the ubx/saturation_* blocks (API change); lower_limits/upper_limits are double arrays of length 1 (broadcast) or data_len.
  • ubx/ramp and ubx/rand are new generic variants; the legacy ubx/ramp_* and ubx/rand_* blocks are retained for backwards compatibility. ramp accumulates in the native type, so 64 bit integer ramps stay exact beyond the 2^53 double mantissa, and rand keeps its PRNG state per instance (erand48/jrand48), making multiple instances independent and thread-safe.

threshold gained an optional hysteresis config (Schmitt trigger): the state switches to 1 only above threshold + hysteresis/2 and back to 0 only below threshold - hysteresis/2, which suppresses event chatter on noisy signals.

tracing: SDT probes and ftrace markers

libubx gained a lightweight, compile-time selectable tracing layer with tracepoints at the framework choke points: step_begin/ step_end around each c-block step and chain_begin/chain_end around each trigger chain run. ptrig additionally emits overrun and dl_overrun events. The backend is picked at build time:

$ cmake -DTRACING=SDT ..     # USDT probes (perf, bpftrace, systemtap)
$ cmake -DTRACING=MARKER ..  # ftrace trace_marker
$ cmake -DTRACING=OFF ..     # default, macros expand to nothing

SDT compiles each probe to a single nop until a consumer attaches, so it can stay enabled in production builds. MARKER writes the events into the kernel ftrace buffer, where they interleave with sched and irq events on a common timeline — which is what finally answers why a cycle overran. This complements the min/max/avg tstats with per-cycle timelines.

rtlog: rebuilt on a lock-free broadcast buffer

The real-time logger was reimplemented on top of the new liblfb, a header-only lock-free broadcast buffer (one producer, any number of independent consumers) that generalises what rtlog used to hand-roll. The shm writer spinlock is gone, replaced by a process-shared, robust priority-inheritance mutex plus C11 atomics for publishing the write offset. This bounds priority inversion between loggers of different priorities, recovers when a writer dies while holding the lock (EOWNERDEAD) instead of deadlocking everyone, and fixes reader-side memory ordering on weakly ordered architectures.

struct ubx_log_msg changed accordingly (ABI change): ts is now an int64_t nanosecond timestamp and level an int32_t, giving an identical, fixed-width layout on 32 and 64 bit targets, and the struct is exactly three cache lines, so log frames never straddle one. UBX_LOG_MSG_MAXLEN shrank from 127 to 115 and became a CMake cache variable. Mixed-version processes must not share a log shm.

ubx-launch: proper option parsing

ubx-launch was ported to optparse: all options gained GNU style long forms (the legacy single-dash long options still work but warn), and -l is now short for --loglevel. New -e/--empty launches an empty node (with stdtypes imported) — useful together with --dbus to load the actual model later via ubx-dbus --load-usc.

ubx-log: syslog forwarding and daemon mode

ubx-log gained two new options: -s tees log messages to syslog (selectable facility via -f LOCAL0..7), and -d daemonizes via libdaemon with a PID file under /run to prevent duplicate instances. Daemon support is now an explicit build option (UBX_LOG_DAEMON). -F disables following, so ubx-log -F dumps the current buffer and exits.

improved enum, struct and union support

Anonymous structs and unions are now handled correctly by cdata.tolua. Named structs and unions gained a custom converter registration mechanism, and enums are properly reflected, with new test coverage throughout.

register (struct) types from Lua

New ubx.type_add(nd, name, cdecl [, doc]) registers a struct type at runtime — no build-time ubx-typegen step or separate C type module needed. It does the ffi.cdef + ffi.sizeof + persistent copy + ubx_type_register in one call, and ubx.type_rm(nd, name) unregisters and frees one it created:

ubx.type_add(nd, "struct point", "struct point { double x; double y; };")
local d = ubx.data_alloc(nd, "struct point")   -- usable like any other type

The registered type is a real, fully-formed ubx_type_t: allocatable, marshallable to/from Lua tables, and visible to tooling. It is node-local (the hash is name-based) — meant for self-contained nodes, not for data crossing a process boundary. This lets a luablock expose a genuine typed config instead of smuggling structured configuration through lua_str. See docs/dev/005. ubx-modinfo show prints the struct definition of such types too.

Lua API docs

ubx.lua is now fully documented with ldoc. The generated API reference is published here and built as part of the standard CMake doc target.

build: switched to CMake

The build system was migrated from make to CMake, reducing boilerplate and improving IDE and cross-compilation support. --version now reports the compile-time options a binary was built with (time source, tracing backend, log message length, ...), and a tidy target runs clang-tidy with the CERT checks.

test cleanup

Tests were reorganised into suites with a single runner, CI was updated to Debian trixie, and the lsdb-intf tests are now part of the standard CI run. The lock-free bits grew multi-writer/multi-reader stress tests for rtlog and SPSC ordering tests for liblfq, both also run under ThreadSanitizer in CI.

various bug fixes

Several minor fixes across ubx.lua, blockdiagram, node cleanup, module reloading, and timing utilities. Notably, blockdiagram re-resolves block references after merging a system, so an overlay can now add connections into the base model instead of having them rejected as unknown blocks.