Although quiet on the master branch, a lot is going on on -dev. Here's a
summary of what's coming (and will become the long awaited v1.0).
new blocks
lsdb-intf and ubx-dbus
A D-Bus interface block that exposes the full node API
(create/remove blocks, connect ports, set/get configs, load USC
models, ...) over D-Bus. The companion ubx-dbus CLI provides
interactive access from the shell.
The plugin extension lets you register custom D-Bus objects at
runtime by loading Lua plugin files, so application-specific
interfaces can sit alongside the standard org.ubx.node API.
The block is based on the
lsdbus bindings. Essentially,
this permits creating coordination interfaces that can cleanly and
safely interact with the hard real-time processing without impacting
real-time performance.
lfrb
A new lock-free ring buffer iblock based on a minimal internal
liblfq implementation which is based on a Vyukov MPMC queue. This
replaces the external liblfds dependency which didn't support
arm64. Strongly typed, hard-real-time safe.
gpio
Linux GPIO block using libgpiod v2. Ports are created dynamically from the
gpios config array — each input line becomes an out-port, each output line
becomes an in-port (type unsigned int). Supports active-low polarity, pull
bias, and emit-on-change mode.
iio
Linux Industrial I/O blocks via libiio for ADCs, DACs, IMUs, and
environmental sensors. Two variants:
ubx/iio— polled, synchronous sysfs read each step; also supports DAC outputubx/iio_buf— buffered, hardware samples atsampling_frequency; ptrig drains the kernel buffer
Ports are created dynamically from the channels config; type is always
double (physical scaled value).
gps
GPS block that reads from gpsd via its shared
memory interface. Emits a single gps out-port of type ubx_gps_data with
position, velocity, fix quality, and uncertainty fields.
webgraph
A luablock that serves an interactive browser graph of the running node
using React Flow + ELK.js. Blocks are shown as nodes (colour-coded by state)
with ports, configs, and iblock edges. The page auto-refreshes and re-runs
the ELK layered layout. Requires luasocket and json.lua; JS libs are
fetched from CDN on first load. Self-trigger it at a low rate alongside your
composition — no separate process needed.
ubx-launch gained a matching -webgraph [PORT] flag to attach one
automatically:
$ ubx-launch -c threshold.usc -webgraph
This replaces the old webif block, which was removed — it had grown too
complex and unmaintainable.
math_double / math_float
Applies any single-argument math.h function element-wise to an input array,
with optional per-element mul and add (y = f(x) * mul + add). The
function is selected via the func config string ("sin", "sqrt", "log",
etc.).
netsink
A luablock-based streaming sink for live-plotting and telemetry
(primarily PlotJuggler). It drains input ports
and emits one message per sample over the network. Two transports —
UDP datagrams or a ZeroMQ PUB socket — and two encodings — JSON
or MessagePack — are selectable, all via lua_str globals. Ports are
created dynamically from the ports declaration:
{ name="sink", type="luablock:netsink" },
-- ...
{ name="sink", config = {
lua_str = [[
ports = { x="double", v="double" }
transport = "zmq" -- "udp" (default) or "zmq"
format = "msgpack" -- "json" (default) or "msgpack"
uri = "tcp://*:9870"
]],
} },
Ports default to scalars; a "[N]" suffix declares an array port
(pos="double[3]"), which serializes as a nested array and is expanded
by PlotJuggler into pos/0, pos/1, ...
Both transports go straight through the LuaJIT FFI — no luasocket or
lzmq dependency. For RT decoupling, run the sink on its own thread
at a lower rate (thread=1, period=100) with deep connection buffers
(buffer_len=N); each step drains the full buffer so no samples are lost
on the RT side.
$ ubx-launch -c netsink.usc
(Previously called udpsink; renamed now that it does more than UDP —
the migration is a one-word type change.)
signal processing: stats, movavg, ewma, mux/demux
A set of small, generic building blocks for composing signal chains. All
are runtime-typed: element type and vector length are chosen via the
type and data_len configs instead of at compile time, so one block
serves all numeric types.
ubx/stats— runningmin,max,mean,stdandcntof the input signal, computed online with Welford's algorithm and emitted as astruct ubx_stat. Vectors are accumulated per element; the output rate can be throttled withstats_output_rate, andskip_firstdiscards N startup samples, so a single transient outlier doesn't ruin min/max and dominate mean and stddev for the rest of the run.ubx/movavg— fixed-window moving average, with amodeconfig selecting the window aggregate (mean,median,minormax). The median is RT-safe (insertion sort into a preallocated scratch buffer).ubx/ewma— exponentially weighted moving average (y += alpha * (x - y)), the cheap, state-free-ish alternative to a window filter.ubx/mux,ubx/demux— composition glue:muxconcatenatesnininput ports into one vector,demuxpartitions a vector ontonoutoutput ports. The optionalin_len/out_lenconfigs assign per-port sub-vector lengths, which makesdemuxdouble as a slicing block — connect only the outputs you need.
vstore and latch
Two new connection iblocks for cases where lfrb's lock-free queue is
paid for but not used. Neither is a default — pick them per connection,
once the respective precondition has been checked.
ubx/vstore is a single-slot value store for a connection whose writer
and reader are stepped by the same trigger in the same thread. That
pair needs neither a queue nor synchronisation: the write always
completes before the read begins. Per connection and step lfrb performs
four atomic read-modify-writes — not uniformly cheap, e.g. on an ARMv8.0
target without LSE each becomes an outline-atomics helper call — and
spreads freeq/usedq/slot/elem over several cache lines, where vstore is
one allocation with the payload next to its flag. It keeps lfrb's read
contract exactly (0 when there is no unread value) and counts violations
of its own precondition on an overwrites port.
ubx/latch is a single-writer, multi-reader latest-value store: a
read does not consume, so any number of readers share one latch and the
value stays available until the writer replaces it. Concurrency is a
seqlock — the writer is wait-free, a reader catching a write in progress
retries and then reports no-data rather than spinning — so unlike vstore
it is safe across threads. Use it for published state sampled by an
unrelated cycle, or for signals only written on change:
{ src="ptrig.overrun_cnt", tgt="overruns.in", type="ubx/latch" },
The catch is the same for both: the return value means a value exists,
not a new value arrived. Anything that must act only on fresh input
stays on lfrb.
other changes
blockdiagram: direct luablock loading
.usc compositions can now instantiate Lua blocks directly using the
luablock:NAME type prefix — no need to manually import the luablock module,
create an instance, and configure lua_file separately:
blocks = {
{ name="ctrl", type="luablock:mycontroller" },
},
The name is resolved against the standard microblx block prefixes.
luablock: self-triggering
luablock now supports built-in self-triggering via the thread and period
configs. For non-RT management or interfacing tasks this eliminates the need for
a dedicated ptrig:
config = { thread=1, period=100 } -- self-trigger at 100 ms
preinit / preexit life-cycle hooks
Blocks gained two optional life-cycle hooks, preinit and preexit,
running before the regular init/cleanup. They add a second
structural-mutation point for the rare case where a block must build its
own interface from its configuration — e.g. create ports whose number or
type depends on a config, or a config whose value drives the creation of
other configs. This is what lets netsink above turn its ports
declaration into real ports before the connections are wired up.
Most blocks never need this — init remains the place for ordinary
setup. Reach for preinit only when interface creation has to depend on
configuration. It's available to C blocks and to luablock (just define
Lua preinit/preexit functions); the rationale is written up in
docs/dev/004.
ptrig: SCHED_DEADLINE support
ptrig now supports Linux EDF scheduling via sched_policy = "SCHED_DEADLINE". Timing parameters are passed via the sched_deadline
config (runtime_ns, deadline_ns, period_ns); deadline_ns and
period_ns default to the period config when 0. After each chain trigger
sched_yield signals budget exhaustion to the kernel. Budget overruns are
caught via SIGXCPU and counted on the deadline_throt_cnt out-port. The
sched_deadline in-port allows runtime parameter updates without
reconfiguration.
ptrig: dynamic period port
ptrig gained a period in-port for runtime adjustment of the trigger period
without reconfiguration.
ptrig: trigger latency measurement
How late does a trigger actually fire? overrun_cnt only counts whole
periods lost, so a trigger that is consistently 200 us late on a 1 ms
period reports zero overruns, and the tstats measure how long the chain
ran, not when it started. The quantity the sleep modes exist to
reduce was the one without a metric — answering it meant reaching for the
kernel's timerlat tracer.
latency_stats = 1 now writes the per-cycle lateness to a new
latency_ns out-port and logs min/max/avg at stop, next to the overrun
report. Connect latency_ns to a ubx/stats block for the stddev. It is
opt-in because it costs one extra clock read per cycle; the first
tstats_skip_first samples are excluded from the logged summary, since
the deadline grid is only anchored on the first cycle and that outlier
dominated the result.
ptrig: hybrid sleep mode and timer slack
sleep_mode gained a third option: hybrid (2) sleeps until
busy_slack_ns before the deadline and busy-waits the rest, which
absorbs the OS wakeup latency at busy-wait accuracy for a
busy_slack_ns / period duty cycle (default 50 us). busy_slack_ns must
exceed the platform's worst-case wakeup latency, otherwise the busy phase
never runs and the mode silently degrades to sleep_mode=0.
The new timerslack_ns config sets the trigger thread's slack via
prctl(PR_SET_TIMERSLACK). Linux defaults non-realtime threads to 50 us
of timer slack, which shows up one-to-one as trigger lateness under
SCHED_OTHER; realtime policies ignore it.
Also fixed: after a missed deadline ptrig discarded the realigned grid point and fired one trigger immediately at an arbitrary phase. It now waits for the realigned deadline, so an overrun drops the affected tick(s) and leaves the phase of all subsequent triggers unchanged.
cheaper clock reads: ubx_gettime_ns
ubx_gettime returns a sec/nsec pair, which on the CNTVCT time source
costs three 64 bit divisions to construct — on a Cortex-A53 the divisions
dominate the call, while the counter read itself is a fraction of it. Most
callers don't want the pair: the busy-wait loops built one per spin
iteration only to compare it, and ptrig converted it straight back.
The new ubx_gettime_ns returns nanoseconds directly; where the counter
frequency divides 1e9 evenly (the common 25/50/100/200 MHz parts) the
conversion is a single multiply, and the odd 19.2/24 MHz parts keep the
division path. Measured on an AM62x at a 500 us period with
sleep_mode=2, mean trigger lateness dropped from 1553 to 1506 ns: the
busy-wait exits closer to the deadline because each spin iteration is
cheaper. This helps the counter-based time sources only — on the
clock_gettime fallback the vDSO call dominates and the arithmetic
around it is noise. Which is worth knowing, because both hardware time
sources still default to off: the out-of-the-box build pays 65 ns per
timestamp against 15 ns for CNTVCT, proportional to how much timing
instrumentation is enabled.
const: runtime in port
ubx/cconst and ubx/iconst gained an in port. Writing to it
updates the held value at runtime — without a reconfigure
cycle. Example use-case: update dynamic configuration safely via the
lsdb-intf block.
runtime-typed saturation, ramp and rand
The compile-time typed block variants are being replaced by single,
runtime-typed blocks that take the element type and vector length from
the type and data_len configs:
ubx/saturationreplaces theubx/saturation_*blocks (API change);lower_limits/upper_limitsaredoublearrays of length 1 (broadcast) ordata_len.ubx/rampandubx/randare new generic variants; the legacyubx/ramp_*andubx/rand_*blocks are retained for backwards compatibility.rampaccumulates in the native type, so 64 bit integer ramps stay exact beyond the 2^53 double mantissa, andrandkeeps its PRNG state per instance (erand48/jrand48), making multiple instances independent and thread-safe.
threshold gained an optional hysteresis config (Schmitt trigger):
the state switches to 1 only above threshold + hysteresis/2 and back
to 0 only below threshold - hysteresis/2, which suppresses event
chatter on noisy signals.
tracing: SDT probes and ftrace markers
libubx gained a lightweight, compile-time selectable tracing
layer with tracepoints at the framework choke points: step_begin/
step_end around each c-block step and chain_begin/chain_end
around each trigger chain run. ptrig additionally emits overrun
and dl_overrun events. The backend is picked at build time:
$ cmake -DTRACING=SDT .. # USDT probes (perf, bpftrace, systemtap)
$ cmake -DTRACING=MARKER .. # ftrace trace_marker
$ cmake -DTRACING=OFF .. # default, macros expand to nothing
SDT compiles each probe to a single nop until a consumer attaches,
so it can stay enabled in production builds. MARKER writes the events
into the kernel ftrace buffer, where they interleave with sched and
irq events on a common timeline — which is what finally answers why
a cycle overran. This complements the min/max/avg tstats with
per-cycle timelines.
rtlog: rebuilt on a lock-free broadcast buffer
The real-time logger was reimplemented on top of the new liblfb,
a header-only lock-free broadcast buffer (one producer, any number of
independent consumers) that generalises what rtlog used to hand-roll.
The shm writer spinlock is gone, replaced by a process-shared, robust
priority-inheritance mutex plus C11 atomics for publishing the write
offset. This bounds priority inversion between loggers of different
priorities, recovers when a writer dies while holding the lock
(EOWNERDEAD) instead of deadlocking everyone, and fixes reader-side
memory ordering on weakly ordered architectures.
struct ubx_log_msg changed accordingly (ABI change): ts is now
an int64_t nanosecond timestamp and level an int32_t, giving an
identical, fixed-width layout on 32 and 64 bit targets, and the struct
is exactly three cache lines, so log frames never straddle one.
UBX_LOG_MSG_MAXLEN shrank from 127 to 115 and became a CMake cache
variable. Mixed-version processes must not share a log shm.
ubx-launch: proper option parsing
ubx-launch was ported to optparse: all options gained GNU style
long forms (the legacy single-dash long options still work but warn),
and -l is now short for --loglevel. New -e/--empty launches an
empty node (with stdtypes imported) — useful together with --dbus
to load the actual model later via ubx-dbus --load-usc.
ubx-log: syslog forwarding and daemon mode
ubx-log gained two new options: -s tees log messages to syslog (selectable
facility via -f LOCAL0..7), and -d daemonizes via libdaemon with a PID
file under /run to prevent duplicate instances. Daemon support is now
an explicit build option (UBX_LOG_DAEMON). -F disables following, so
ubx-log -F dumps the current buffer and exits.
improved enum, struct and union support
Anonymous structs and unions are now handled correctly by cdata.tolua. Named
structs and unions gained a custom converter registration mechanism, and enums
are properly reflected, with new test coverage throughout.
register (struct) types from Lua
New ubx.type_add(nd, name, cdecl [, doc]) registers a struct type at
runtime — no build-time ubx-typegen step or separate C type module
needed. It does the ffi.cdef + ffi.sizeof + persistent copy +
ubx_type_register in one call, and ubx.type_rm(nd, name) unregisters
and frees one it created:
ubx.type_add(nd, "struct point", "struct point { double x; double y; };")
local d = ubx.data_alloc(nd, "struct point") -- usable like any other type
The registered type is a real, fully-formed ubx_type_t: allocatable,
marshallable to/from Lua tables, and visible to tooling. It is node-local
(the hash is name-based) — meant for self-contained nodes, not for data
crossing a process boundary. This lets a luablock expose a genuine typed
config instead of smuggling structured configuration through lua_str. See
docs/dev/005.
ubx-modinfo show prints the struct definition of such types too.
Lua API docs
ubx.lua is now fully documented with
ldoc. The generated API
reference is published here and
built as part of the standard CMake doc target.
build: switched to CMake
The build system was migrated from make to CMake, reducing boilerplate and
improving IDE and cross-compilation support. --version now reports the
compile-time options a binary was built with (time source, tracing backend,
log message length, ...), and a tidy target runs clang-tidy with the CERT
checks.
test cleanup
Tests were reorganised into suites with a single runner, CI was updated to
Debian trixie, and the lsdb-intf tests are now part of the standard CI run.
The lock-free bits grew multi-writer/multi-reader stress tests for rtlog
and SPSC ordering tests for liblfq, both also run under ThreadSanitizer
in CI.
various bug fixes
Several minor fixes across ubx.lua, blockdiagram, node cleanup, module
reloading, and timing utilities. Notably, blockdiagram re-resolves block
references after merging a system, so an overlay can now add connections
into the base model instead of having them rejected as unknown blocks.

