Skip to content

XAP — Xylolabs Audio Protocol Specification

PATENT PENDING — XAP (Xylolabs Audio Protocol) is a proprietary technology of Xylolabs Inc. Patent applications have been filed. Unauthorized use, reproduction, or distribution is prohibited.

Revision: 2026-07-08


Table of Contents

  1. Overview
  2. Container Formats
  3. Configuration
  4. Platform Requirements
  5. Platform Compatibility Matrix
  6. Audio Quality Characteristics
  7. SDK Integration
  8. Comparison with IMA-ADPCM
  9. Related Documents
  10. Legal

1. Overview

XAP (Xylolabs Audio Protocol) is Xylolabs' proprietary MDCT-based spectral audio codec designed for real-time compression of multi-channel audio on resource-constrained IoT and industrial monitoring hardware.

Key Properties

Property Value
Transform Modified Discrete Cosine Transform (MDCT)
Sample rates 8, 16, 24, 32, 48, 96 kHz
Channels 1–4 (encoded independently)
Frame durations 7.5 ms, 10 ms
Compression ratio 8:1–10:1 (bitrate-dependent)
Bitrate range 16–320 kbps per channel
Typical bitrate 64–80 kbps per channel
Algorithmic delay 7.5–10 ms (one frame)
CPU requirement ~10 MIPS per channel (with DSP acceleration)
RAM per channel ~8 KB encoder state
Codec ID (XMBP) 0x03

These six sample rates are served by two distinct on-wire container formats — see §2 Container Formats for the byte-level layouts and the rate-based dispatch rule (§2.3).

Design Goals

XAP is designed for industrial audio monitoring applications with the following constraints:

  • MCU encode, server decode: The encoder runs on bare-metal or RTOS-based MCUs with no operating system. The decoder runs on the server with no resource constraints.
  • LTE-M1 bandwidth budget: Four channels at 96 kHz must fit within approximately 30–40 KB/s sustained uplink, achieved at 64–80 kbps per channel (32–40 KB/s total).
  • No dynamic memory allocation: The encoder uses only statically allocated buffers. All state fits within a single XapEncoder struct.
  • FPU/DSP acceleration: XAP requires a hardware floating-point unit and benefits significantly from ARM DSP extensions or SIMD instruction sets. Platforms lacking an FPU must use IMA-ADPCM instead.

2. Container Formats

"XAP" on the wire is actually two distinct container formats, selected by the encoder based on sample rate:

  • §2.1 Standard LC3 transport — the default path for 16, 24, 32, and 48 kHz. Wraps standards-aligned LC3 (Low Complexity Communication Codec) frames — produced by the lc3-codec crate — in a lightweight framing compatible with FFmpeg's .lc3 container.
  • §2.2 Legacy XAP MDCT — Xylolabs' original proprietary MDCT codec, used for 8 kHz and 96 kHz, where the standards LC3 backend is unavailable (§2.3).

Both formats travel the same .xap/.lc3 upload path and share codec identifier 0x03 in XMBP; a receiver disambiguates them at decode time using the tag check in §2.3.

2.1 Standard LC3 Transport (Default, 16–48 kHz)

For 16, 24, 32, and 48 kHz sample rates, the SDK encoder (crates/xylolabs-sdk/src/codec/lc3.rs) emits standards-aligned LC3 audio frames wrapped in a transport container compatible with FFmpeg's .lc3 muxer. The server decoder (crates/xylolabs-transcode/src/xap_decode.rs) implements the inverse.

2.1.1 Transport Header

Fixed 18-byte header (20 bytes when the optional HR-mode field is present). The leading tag is big-endian; every other multi-byte field is little-endian:

Offset Size Field Encoding Description
0 2 tag u16 BE Transport magic, always 0x1CCC
2 2 header_size u16 LE Header length in bytes: 18 normally, 20 when hr_mode is present
4 2 sample_rate u16 LE Sample rate in Hz ÷ 100 (e.g. 480 = 48,000 Hz)
6 2 bitrate u16 LE Aggregate bitrate across all channels, in bps ÷ 100 — not per-channel
8 2 channels u16 LE Channel count, validated to 1–4
10 2 frame_us u16 LE Frame duration in µs ÷ 10 (e.g. 1000 = 10 ms)
12 2 ep_mode u16 LE Error-protection mode; nonzero = enabled. The SDK encoder always writes 0
14 4 total_samples_per_channel u32 LE PCM sample count per channel carried by this report
18 2 hr_mode u16 LE Optional — present only when header_size >= 20. High-resolution mode flag, nonzero = enabled. Never emitted by the SDK encoder (header_size is always 18); parsed defensively for interop with other .lc3 producers
Offset:   0     2       4       6       8       10      12      14              18     [20]
        +-----+-------+-------+-------+-------+-------+-------+---------------+-------+
        | tag | hdrsz | srate | brate | chans | fr_us | ep_md | total_samples | hr_md |
        |u16BE|u16 LE |u16 LE |u16 LE |u16 LE |u16 LE |u16 LE |    u32 LE     |u16 LE |
        +-----+-------+-------+-------+-------+-------+-------+---------------+-------+
                                                                                (optional)

Header-level validation is broader than what is actually produced or decoded. parse_lc3_transport_header accepts sample_rate values of 8,000 / 16,000 / 24,000 / 32,000 / 48,000 / 96,000 Hz and frame_us values of 2,500 / 5,000 / 7,500 / 10,000 µs at the header-parsing stage — a superset of what the SDK encoder ever emits (16/24/32/48 kHz, 7.5/10 ms only; see §2.3) and of what the bundled decoder can fully decode (§2.1.4). This exists for interop tolerance with third-party .lc3 producers, not because the Xylolabs pipeline itself generates those combinations.

2.1.2 Packet Stream

Immediately following the header, the payload is a sequence of [packet_len: u16 LE][packet: packet_len bytes] units, one per encoded frame period, repeated until the data ends:

[packet_len: u16 LE][packet bytes]   <-- repeated
                     |
                     +-- channel_0 block: packet_len / channels bytes
                         channel_1 block: packet_len / channels bytes
                         ...
                         channel_{N-1} block: packet_len / channels bytes

Each packet holds the LC3-encoded bytes for all channels of one frame, laid out as channels fixed-size blocks concatenated in channel order — channel 0's bytes, then channel 1's, and so on. Samples are not byte-interleaved across channels within a packet. bytes_per_channel = packet_len / channels; a packet whose length does not divide evenly by the header's channel count ends the decodable stream at that point (§2.1.4).

Frame samples per channel are derived from the header rather than stored per-packet: frame_samples = sample_rate_hz * frame_us / 1_000_000 (e.g. 48 kHz @ 10 ms → 480 samples/channel). The decoder tracks a running remaining_samples counter, seeded from the header's total_samples_per_channel and decremented by min(remaining_samples, frame_samples) per packet, so a final short packet is trimmed rather than zero-padded into the output.

2.1.3 Multi-Report Resync (cycle 4, AGG-C4-17)

A single stored or streamed buffer may contain multiple back-to-back "reports" — for example, an encoder restart or a device reconnect mid-recording. Each report is a complete transport header followed by its own packet stream. At every packet-boundary read, the decoder re-checks whether the next 2 bytes, read as big-endian, equal the 0x1CCC tag:

  • Tag matches, and the new header's channels / sample_rate_hz / frame_us match the report already in progress: resync in place. The new header's total_samples_per_channel is added to remaining_samples, packet_offset advances by the new header's header_size, and decoding continues as one continuous stream with no gap in the output.
  • Tag matches, but a parameter differs: decoding stops; only the prefix decoded so far is kept. A parameter change would require a fresh decoder instance and is not concatenated.
  • No tag match: the 2 bytes are read as an ordinary packet_len and decoding proceeds per §2.1.2.

This disambiguation is unambiguous in practice: a colliding packet_len would require a real packet of exactly 52,252 bytes (0xCC1C read little-endian — the same two bytes that read as 0x1CCC big-endian), far beyond any packet the encoder produces (packet size is bounded by configured per-channel bitrate × frame duration; even at the highest supported bitrate this stays well under a few KB).

2.1.4 Truncation and Error Handling

  • Zero-length or overrunning packet_len: treated as a truncated tail (device disconnect, link drop, mid-frame cut) — decoding stops at the last clean packet boundary and the decoded prefix is kept, rather than discarding the whole recording.
  • Packet length not evenly divisible by channels: same treatment — ends the decodable prefix.
  • remaining_samples reaches 0 with no further same-parameter report following: decoding stops normally.
  • Sample rate or frame duration outside decoder support: a header that passes the broader header-level whitelist (§2.1.1) but falls outside {8000, 16000, 24000, 32000, 48000} Hz or {7500, 10000} µs fails at decode time with an explicit "not supported by the bundled decoder" error (this excludes 96 kHz and the 2.5/5 ms durations specifically). In practice the SDK encoder never emits a standard-LC3 header outside 16/24/32/48 kHz × 7.5/10 ms (§2.3), so this case only matters for third-party .lc3 producers.

2.2 Legacy XAP MDCT (8 kHz and 96 kHz)

Scope note. The MDCT algorithm and wire format below are inherently rate-agnostic — they were originally designed to cover all six XAP sample rates (8–96 kHz) and remain available at any of them via the legacy::XapEncoder type (embedded firmware may construct it directly, e.g. for a PSRAM-backed cosine table at 8 kHz). As of the encoder rewrite that added §2.1, however, the SDK's default XapEncoder::new() constructor only automatically routes 8 kHz and 96 kHz through this path (§2.3) — 16/24/32/48 kHz now default to Standard LC3 transport. Tables in this section and in §3 Configuration that show entries for 16/24/32/48 kHz describe what this algorithm does at those rates when explicitly constructed via legacy::XapEncoder, not what the default SDK encoder emits on the wire today.

2.2.1 Signal Flow

Interleaved PCM input (i16[])
         |
         v
   De-interleave
   (per-channel extraction)
         |
         v
  MDCT Forward Transform
  N samples --> N/2 spectral coefficients
         |
         v
  Adaptive Quantization
  (step size derived from mean absolute coefficient)
         |
         v
  Coefficient Packing
  (2-byte quant header + 8-bit quantized coefficients)
         |
         v
  XAP Frame Output
  (5-byte frame header + per-channel payloads)

Each channel is encoded independently. The encoder accepts interleaved multi-channel PCM input, extracts each channel's samples into a contiguous mono buffer, applies the MDCT, quantizes the resulting spectral coefficients, and packs them into the output frame.

2.2.2 MDCT Forward Transform

The Modified Discrete Cosine Transform converts N time-domain PCM samples into N/2 frequency-domain spectral coefficients. The basis function is:

X[k] = sum_{n=0}^{N-1} x[n] * cos(pi/N * (n + 0.5 + N/4) * (k + 0.5))

  where:
    N   = frame_samples (number of PCM samples per channel per frame)
    k   = coefficient index, 0 <= k < N/2
    n   = sample index, 0 <= n < N

The transform produces N/2 real-valued coefficients that represent the spectral energy of the frame.

Precomputed Cosine Table

For frame sizes where N <= 320 (sample rates up to 32 kHz at 10 ms, or up to 24 kHz at 7.5 ms), the encoder precomputes a fixed-point cosine table at initialization time:

cos_table[k * N + n] = (cos(pi/N * (n + 0.5 + N/4) * (k + 0.5)) * 32768) as i32

Table dimensions: (N/2) x N entries of i32. Maximum size: 320 samples → 51,200 entries → 200 KB.

Using the precomputed table, the MDCT inner loop reduces to fixed-point multiply-accumulate with no trigonometric function calls, yielding O(N²) multiplies at very low constant overhead. Measured encode times on M-series host:

Sample Rate Frame Samples Table Used Encode Time (1ch, 10ms)
8 kHz 80 Yes 0.5 µs
16 kHz 160 Yes 1.1 µs
24 kHz 240 Yes 1.9 µs
32 kHz 320 Yes (limit) 3.0 µs
48 kHz 480 No 512.3 µs
96 kHz 960 No 1954.0 µs

The 170x discontinuity at 48 kHz (software fallback path) is eliminated on MCU targets by DSP-accelerated MDCT paths (see §4 Platform Requirements).

Stack Allocation Note

XapEncoder contains a 200 KB cosine table. On embedded targets, allocate the encoder in a static variable rather than on the stack. The table is only populated when frame_samples <= 320.

2.2.3 Adaptive Quantization

After the MDCT, each coefficient X[k] is quantized to a signed 8-bit integer using an adaptive step size:

Step size derivation:

mean_abs = sum(|X[k]|) / (N/2)
step     = clamp(mean_abs / 64, 1, 8192)

The step size adapts to the signal level of each frame, distributing the available quantization range across the actual coefficient magnitudes. A larger step size is used for louder frames; a smaller step size for quieter frames.

Quantization:

Q[k] = clamp(X[k] / step, -128, 127)

Coefficients that exceed the representable range after quantization are clipped to [-128, 127].

2.2.4 Per-Channel Budget Allocation

The total frame payload is divided equally among channels after subtracting the 5-byte frame header:

payload_bytes  = frame_bytes - 5
per_ch_budget  = max(payload_bytes / channels, 4)

Each channel's packed output is zero-padded to exactly per_ch_budget bytes, producing fixed-size frames suitable for streaming without framing overhead.

2.2.5 Wire Format

2.2.5.1 Frame Header (5 bytes)

Every legacy XAP frame begins with a 5-byte fixed header in big-endian byte order:

Offset  Size  Field         Type    Description
------  ----  -----------   ------  ------------------------------------------
0       2     frame_samples u16 BE  PCM samples per channel in this frame
2       1     channels      u8      Number of channels (1–4)
3       2     frame_bytes   u16 BE  Total frame size including this header
 0               1               2               3               4
 7 6 5 4 3 2 1 0 7 6 5 4 3 2 1 0 7 6 5 4 3 2 1 0 7 6 5 4 3 2 1 0 7 6 5 4 3 2 1 0
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|        frame_samples (u16 BE)         |   channels    |      frame_bytes (u16 BE)
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
                                                         ^-- continues byte 3..4 -->

This 5-byte header is also what the format-dispatch tag check (§2.3) reads when a stream does not begin with 0x1CCC: frame_samples occupies the same leading 2 bytes as the standard transport's tag field, and is guaranteed by construction to never alias it (§2.3).

2.2.5.2 Per-Channel Payload

Immediately following the frame header, the payload contains channels sequential channel blocks, each of exactly per_ch_budget bytes. Within each channel block:

Offset  Size  Field         Type    Description
------  ----  -----------   ------  ------------------------------------------
0       2     step_size     u16 BE  Quantization step size for this channel
2       N     coeffs[]      i8[]    Quantized MDCT coefficients (up to per_ch_budget - 2)

The number of coefficient bytes packed is min(N/2, per_ch_budget - 2). If the packed coefficient count is less than per_ch_budget - 2, the remaining bytes are zero-padded to maintain the fixed frame size.

2.2.5.3 Complete Frame Structure
[frame_samples: 2B][channels: 1B][frame_bytes: 2B]   <- 5-byte header
[step_ch0: 2B][q0[0], q0[1], ..., q0[M-1]]           <- channel 0 block (per_ch_budget bytes)
[step_ch1: 2B][q1[0], q1[1], ..., q1[M-1]]           <- channel 1 block (per_ch_budget bytes)
...
[step_chN: 2B][qN[0], qN[1], ..., qN[M-1]]           <- channel N block (per_ch_budget bytes)

  where M = per_ch_budget - 2 (coefficient bytes per channel)

Legacy frames concatenate directly one after another with no outer container — the transcode pipeline's legacy decoder (§7.3) walks the buffer by reading each frame's own frame_bytes field to find the next frame's start.

2.3 Format Dispatch Rule

How a decoder (or any other consumer) determines which container format a given .xap/.lc3 byte buffer uses:

  1. Tag check (primary, byte-level). Read the first 2 bytes as big-endian u16. If they equal 0x1CCC, the stream is Standard LC3 transport (§2.1); parse the 18(+2)-byte header immediately following. Otherwise, treat the stream as Legacy XAP MDCT (§2.2) and parse its first 5 bytes as a legacy frame header. This same check is repeated at each of the decoder library's entry points.
  2. The check is unambiguous by construction: a legacy frame's first 2 bytes are its frame_samples field, one of a fixed set of values ≤ 960 (§2.2.4 / §3.2 table), and 0x1CCC = 7,372 is asserted at compile time to exceed that maximum — assert!(LC3_TRANSPORT_TAG > 960) — so no valid legacy frame can be misread as the standard-transport tag.
  3. Sample-rate → path mapping (encoder-side). The SDK encoder (XapEncoder::new) selects the container by configured sample rate before any bytes are written:
Sample Rate Container Emitted Reason
16, 24, 32, 48 kHz Standard LC3 transport Default; the standards LC3 backend supports these rates
8 kHz Legacy XAP MDCT Upstream lc3-codec panics during bandwidth detection at 8 kHz
96 kHz Legacy XAP MDCT Standards LC3 backend does not support 96 kHz frames
  1. Codec ID in XMBP. When carried inside an XMBP (Xylolabs Metadata Batch Protocol) envelope, both XAP containers share codec identifier 0x03 — XMBP itself does not distinguish standard-LC3 from legacy-MDCT; the tag check in step 1 is what a receiver uses to pick a decode path once it has the frame bytes in hand. See XMBP-SPECIFICATION.md for the full wire format of XMBP envelopes.

3. Configuration

3.1 XapConfig Fields

The XapConfig struct controls encoder behavior for both container formats — it is the single entry point (XapEncoder::new) that dispatches per §2.3:

pub struct XapConfig {
    /// Sample rate in Hz.
    /// Supported: 8000, 16000, 24000, 32000, 48000, 96000
    pub sample_rate: u32,

    /// Frame duration in microseconds.
    /// Supported: 7500 (7.5 ms) or 10000 (10 ms)
    pub frame_duration: u32,

    /// Target bitrate in bits/sec per channel.
    /// 0 = auto (~8:1 compression ratio).
    /// Valid range when non-zero: 16000–320000 bps
    pub bitrate: u32,

    /// Number of audio channels. Range: 1–4.
    pub channels: u8,
}

3.2 Supported Sample Rates

Sample Rate Frame Samples (7.5 ms) Frame Samples (10 ms) Cosine Table Used
8,000 Hz 60 80 Yes
16,000 Hz 120 160 Yes
24,000 Hz 180 240 Yes
32,000 Hz 240 320 Yes (limit)
44,100 Hz Not supported
48,000 Hz 360 480 No (runtime cosf)
96,000 Hz 720 960 No (runtime cosf)

44.1 kHz is not supported. All other standard professional rates up to 96 kHz are supported.

The "Cosine Table Used" column describes the legacy MDCT algorithm's behavior at each rate (§2.2). Per the dispatch rule in §2.3, only the 8,000 Hz and 96,000 Hz rows are reachable via XapEncoder::new()'s default routing today; 16,000–48,000 Hz route through Standard LC3 transport (§2.1), which has no cosine-table concept.

3.3 Bitrate and Compression

Bitrate (per channel) Compression Ratio Typical Use
16–32 kbps ~12:1–24:1 Highly constrained bandwidth
64 kbps ~10:1 Standard industrial monitoring (recommended)
80 kbps ~8:1 Broadcast-grade quality
128 kbps ~5:1 High-fidelity monitoring
320 kbps ~2:1 Near-lossless

When bitrate = 0, the encoder targets approximately 8:1 compression based on the PCM frame size.

Frame byte calculation from bitrate:

frame_bytes = bitrate * frame_duration_us / 8_000_000
              clamped to [20, 1200]

3.4 Default Configuration

The SDK default for the XapEncoder::new_default(sample_rate) constructor:

XapConfig {
    sample_rate:    <caller-supplied>,
    frame_duration: 10_000,   // 10 ms
    bitrate:        64_000,   // 64 kbps per channel
    channels:       1,
}

The SDK-level Config default uses audio_sample_rate = 16_000, audio_channels = 4, audio_batch_ms = 500 — i.e. the SDK default routes through Standard LC3 transport (§2.1) unless overridden.


4. Platform Requirements

4.1 Mandatory: Hardware FPU

XAP requires a hardware floating-point unit. The MDCT computation uses single-precision floating-point arithmetic (f32). Platforms without an FPU must use IMA-ADPCM.

Minimum supported core: ARM Cortex-M4F or equivalent (single-precision FPU + DSP extensions).

Platforms without FPU (Cortex-M0+, Cortex-M3, RISC-V without F extension) cannot run XAP. Use features = ["adpcm"] on these targets.

DSP extensions (ARM SMLAD/SMLAL saturating MAC, or Xtensa PIE SIMD) dramatically reduce MDCT compute cost:

DSP Architecture Platforms Mechanism XAP Speedup
ARMv8-M DSP (Cortex-M33) RP2350, nRF9160 Dual 16x16 MAC (SMLAD), saturating arithmetic ~30%
Cortex-M4F FPU+DSP STM32WB55 Hardware float MAC + arm_rfft_fast_f32 (3–5x for MDCT) ~35–40%
Xtensa PIE SIMD (ESP32-S3) ESP32-S3 128-bit SIMD: 4x f32 or 8x i16 per instruction ~60%

4.3 Resource Table

Encoder state and buffer requirements for common configurations (per encoder instance):

Configuration XAP Encoder State Cosine Table Ring Buffer XMBP Frame Total
1ch @16kHz, 10ms ~1 KB 10 KB 4 KB 2 KB ~17 KB
4ch @16kHz, 10ms ~4 KB 10 KB 16 KB 4 KB ~34 KB
4ch @48kHz, 10ms ~4 KB 32 KB 8 KB ~44 KB
4ch @96kHz, 10ms ~4 KB 64 KB 8 KB ~76 KB

The cosine table (200 KB maximum for N <= 320) is embedded within the XapEncoder struct. For 48 kHz and above (N > 320), no cosine table is allocated; the encoder falls back to runtime computation on the host, or uses DSP-accelerated FFT on MCU targets.

4.4 DSP Acceleration Details

ARM Cortex-M33/M4F — CMSIS-DSP

Enable with Cargo feature cmsis-dsp or C define XYLOLABS_USE_CMSIS_DSP=1:

  • The MDCT forward path dispatches to a dual-MAC inner loop using SMLAD (Cortex-M33 fixed-point) or arm_rfft_fast_f32 (Cortex-M4F floating-point).
  • Processes two samples per iteration using dual 16x16 multiply-accumulate.
  • Saturating arithmetic (QADD, SSAT) eliminates branch-based coefficient clipping.
  • Speedup: 30–40% reduction in MDCT compute time (30–60 MIPS savings at 4ch @96kHz).

Xtensa LX7 — ESP32-S3 PIE SIMD

Enable with Cargo feature esp32-simd or C define XYLOLABS_USE_ESP32S3_SIMD=1:

  • The MDCT inner loop processes 4 samples per iteration using 128-bit vector registers.
  • Maps to ESP32-S3 PIE (Processor Instruction Extensions) vector operations on real hardware.
  • Hardware AES/SHA offloads TLS from the main CPU, freeing additional cycles for codec.
  • Speedup: up to 60% reduction in total XAP CPU usage (50 → 20 MIPS for MDCT at 4ch @96kHz).

5. Platform Compatibility Matrix

Evaluated using measured MIPS profiles from burn-in testing and platform datasheets. CPU% represents total system utilization (codec + I/O + transport + sensors + housekeeping) for the maximum supported audio configuration.

Target Core Clock SRAM DSP/FPU Max Audio Config CPU% RAM% Verdict
RP2350 (Pico 2) Cortex-M33 150 MHz 520 KB M33 DSP+FPU 4ch @96kHz XAP 46.0% 16.9% COMFORTABLE
ESP32-S3 Xtensa LX7 240 MHz 512 KB + 8 MB PSRAM PIE SIMD+FPU 4ch @96kHz XAP 17.7% 24.2% COMFORTABLE
nRF9160 Cortex-M33 64 MHz 256 KB M33 DSP+FPU 2ch @48kHz XAP 44.5% 21.9% FEASIBLE
STM32WB55 Cortex-M4F 64 MHz 256 KB M4F DSP+FPU 2ch @48kHz XAP 42.2% 17.2% FEASIBLE

Verdict definitions:

Verdict CPU Utilization Meaning
COMFORTABLE < 50% Ample headroom for OTA updates, additional processing, or future features.
FEASIBLE 50–70% Sufficient for stable operation with careful task scheduling.
TIGHT 70–85% Operational but may exhibit jitter under worst-case interrupt latency.
MARGINAL 85–100% Risk of frame drops under load. Not recommended for production.
ADPCM ONLY N/A No DSP/FPU; XAP encoder cannot run. IMA-ADPCM at 4:1 compression only.
SENSOR ONLY N/A Extreme SRAM constraint. Minimal ADPCM (1–2ch) plus sensor telemetry only.

5.1 Detailed CPU Budget — RP2350 (4ch XAP @96kHz)

Dual-core Cortex-M33 at 150 MHz. Core 0: I2S DMA + XAP encode. Core 1: XMBP + HTTP + sensors.

Component Baseline MIPS With DSP MIPS % of 150 MHz
I2S DMA handling 2 2 1.3%
XAP MDCT forward 50 35 23.3%
XAP quantize+pack 15 10 6.7%
XMBP batch encode 5 5 3.3%
HTTP transport 10 10 6.7%
Sensor sampling (26ch) 5 5 3.3%
Watchdog + housekeeping 2 2 1.3%
Total 89 69 46.0%
Available headroom 61 81 54.0%

5.2 Detailed CPU Budget — ESP32-S3 (4ch XAP @96kHz)

Dual-core Xtensa LX7 at 240 MHz (480 MIPS total).

Component Baseline MIPS With PIE MIPS % of 480 MHz
I2S DMA handling 2 2 0.4%
XAP MDCT forward 50 20 4.2%
XAP quantize+pack 15 6 1.3%
WiFi stack (FreeRTOS) 30 30 6.3%
XMBP batch encode 5 5 1.0%
HTTP/TLS transport 20 12 2.5%
Sensor sampling (26ch) 5 5 1.0%
PSRAM DMA management 3 3 0.6%
Watchdog + housekeeping 2 2 0.4%
Total 132 85 17.7%
Available headroom 348 395 82.3%

6. Audio Quality Characteristics

The characteristics in this section describe the legacy MDCT algorithm's custom 8-bit adaptive quantizer (§2.2.3) specifically. Standard LC3 transport (§2.1) uses the published LC3 standard's own quantization and psychoacoustic model rather than XAP's custom quantizer, so its quality-per-bitrate profile is not necessarily the same as what is tabulated below. Re-benchmarking the Standard LC3 path against these tables is out of scope for this revision.

6.1 Perceptual Quality by Bitrate

Bitrate (per channel) Quality Level Description
16–32 kbps Basic Recognizable audio; significant spectral artifacts at high frequencies. Suitable for voice monitoring.
48 kbps Good Acceptable for broadband monitoring; minor artifacts above 12 kHz.
64 kbps Near-transparent Perceptually transparent for most industrial monitoring applications. Recommended default.
80 kbps Broadcast-grade Indistinguishable from original in blind tests. Full-spectrum fidelity to 48 kHz.
128+ kbps High-fidelity Reference quality. Full spectral content preserved.

6.2 Latency

XAP introduces exactly one frame of algorithmic delay:

Frame Duration Algorithmic Delay
7.5 ms 7.5 ms
10 ms 10 ms

There is no look-ahead or overlap buffer in the encoder. End-to-end latency is bounded by frame duration plus network transit time. For real-time monitoring dashboards, the 10 ms frame duration is preferred as the longer encode window provides more headroom for MCU interrupt jitter.

6.3 Channel Scaling

XAP scales sub-linearly with channel count due to amortized per-frame overhead:

Channels @16kHz Encode Time (host) Per-Channel Scaling Factor
1 1.0 µs 1.0 µs 1.00x
2 1.9 µs 1.0 µs 1.90x
3 2.8 µs 0.9 µs 2.80x
4 3.7 µs 0.9 µs 3.70x

The ~8% sub-linear efficiency gain per channel comes from amortized frame header writes and de-interleave overhead. Per-channel MDCT cost dominates.


7. SDK Integration

7.1 Rust SDK

Add the xap feature to the xylolabs-sdk dependency in Cargo.toml. For DSP-accelerated targets, also enable the appropriate DSP feature:

# RP2350 / nRF9160 (Cortex-M33) — CMSIS-DSP fixed-point path
xylolabs-sdk = { path = "../../crates/xylolabs-sdk", features = ["xap", "cmsis-dsp"] }

# STM32WB55 (Cortex-M4F) — CMSIS-DSP floating-point path
xylolabs-sdk = { path = "../../crates/xylolabs-sdk", features = ["xap", "cmsis-dsp"] }

# ESP32-S3 (Xtensa LX7 with PIE) — PIE SIMD floating-point path
xylolabs-sdk = { path = "../../crates/xylolabs-sdk", features = ["xap", "esp32-simd"] }

# Platforms without DSP/FPU — ADPCM only
xylolabs-sdk = { path = "../../crates/xylolabs-sdk", default-features = false, features = ["adpcm"] }

Encoder usage:

use xylolabs_sdk::codec::lc3::{XapEncoder, XapConfig};

let encoder = XapEncoder::new(XapConfig {
    sample_rate:    16_000,
    frame_duration: 10_000,   // 10 ms
    bitrate:        64_000,   // 64 kbps
    channels:       4,
}).expect("invalid XAP configuration");

// Encode one frame of interleaved PCM
let mut out = vec![0u8; encoder.frame_bytes() as usize];
let written = encoder.encode_frame(&pcm_samples, &mut out);

XapEncoder::new dispatches to Standard LC3 transport or Legacy XAP MDCT per §2.3 based on sample_rate — the call site above (16 kHz) produces Standard LC3 transport framing (§2.1).

7.2 C SDK

Use the compile-time defines in config.h to select the codec and DSP path. DSP acceleration is auto-detected from compiler flags:

#define XYLOLABS_CODEC_XAP       1    /* enable XAP encoder */
#define XYLOLABS_USE_CMSIS_DSP   1    /* Cortex-M4F / M33 targets */
/* or */
#define XYLOLABS_USE_ESP32S3_SIMD 1   /* ESP32-S3 targets */

Override explicitly via CMake when auto-detection is insufficient:

target_compile_definitions(my_firmware PRIVATE
    XYLOLABS_CODEC_XAP=1
    XYLOLABS_USE_CMSIS_DSP=1
)

7.3 Server-Side Decoder

The Xylolabs server includes a decoder library (crates/xylolabs-transcode/src/xap_decode.rs) that reconstructs PCM from either container format (§2). It dispatches per §2.3, then follows one of two decode paths:

Standard LC3 transport (§2.1):

  1. Check the leading 2 bytes for the 0x1CCC tag (big-endian).
  2. Parse the 18(+2)-byte transport header (§2.1.1): sample rate, channel count, frame duration, total sample count.
  3. Walk the [packet_len][packet] stream (§2.1.2), splitting each packet into its per-channel blocks.
  4. Decode each per-channel block with the bundled standards LC3 decoder (lc3_codec::decoder::lc3_decoder::Lc3Decoder) — a standards-defined inverse transform plus residual/LTPF reconstruction, not XAP-specific.
  5. Resync across embedded headers for multi-report streams (§2.1.3); stop at truncated tails (§2.1.4).
  6. Interleave the decoded per-channel PCM.

Legacy XAP MDCT (§2.2):

  1. Parse frame header (5 bytes): extract frame_samples, channels, frame_bytes.
  2. Per-channel dequantization: coeff_f32 = quantized_i8 * step_size.
  3. Inverse MDCT: x[n] = (2/N) * Σ X[k] * cos(π/N * (n + 0.5 + N/4) * (k + 0.5)).
  4. Clamp and convert to i16, interleave channels.

Both paths read frame/stream parameters directly from their own header, making either format self-describing without an external content-type. Sample rate for the legacy path is inferred from frame_samples (e.g. 160 samples → 16 kHz at 10 ms); the standard path reads sample rate directly from the transport header.

Transcode pipeline integration: When a .xap/.lc3 file is uploaded via /api/v1/uploads, the transcode pipeline detects the container per §2.3, decodes it to a temporary PCM WAV file, then passes it to FFmpeg for transcoding to the target format (Opus, FLAC, MP3, etc.).


8. Comparison with IMA-ADPCM

XAP and IMA-ADPCM are the two codecs supported by the Xylolabs SDK. This table summarizes the trade-offs:

As in §6, the XAP-side figures below primarily characterize the legacy MDCT path (§2.2); Standard LC3 transport (§2.1) is a different, standards-defined codec and its quality/CPU/RAM profile may differ from the legacy numbers shown here.

Property XAP IMA-ADPCM
Algorithm MDCT spectral transform Sample-by-sample delta quantization
Compression ratio 8:1–10:1 4:1 (fixed)
Bitrate range 16–320 kbps/ch Fixed at sample_rate × 0.5 bytes/s
Perceptual quality @64kbps Near-transparent Fair — audible quantization noise
Spectral fidelity Full spectrum preserved High-frequency rolloff under load
Algorithmic delay 7.5–10 ms (1 frame) Sample-level (< 0.1 ms effective)
CPU (with DSP) ~10 MIPS/ch < 1 MIPS/ch
CPU (without DSP) Not feasible < 1 MIPS/ch
RAM per channel ~8 KB < 1 KB
FPU required Yes No
DSP recommended Yes No
Platforms supported M4F, M33, Xtensa LX7 All platforms
Implementation complexity Higher (MDCT + quant) Low (lookup tables only)
Typical use case Industrial monitoring, full-spectrum audio Legacy sensors, CPU-constrained MCUs
Integer-only arithmetic No (requires FPU) Yes (pure integer)

Selection guidance:

  • Use XAP when the MCU has an FPU (Cortex-M4F or better) and bandwidth efficiency matters. XAP achieves the same audio quality at half the bandwidth of ADPCM.
  • Use IMA-ADPCM on platforms without FPU, or when CPU budget for audio is less than 2 MIPS per channel, or when per-channel RAM must be below 2 KB.
  • For mixed deployments, the XMBP protocol supports both codecs on the same stream endpoint. The server decoder auto-selects based on the codec identifier in the XMBP stream header.

Document Description
XMBP-SPECIFICATION.md Xylolabs Metadata Batch Protocol wire format; XAP frames are carried as XMBP audio stream payloads
CODEC-ANALYSIS.md Comparative analysis of 16+ audio codecs across 5 MCU platforms; detailed justification for XAP selection
PERFORMANCE-EVALUATION.md Benchmark data: encode times per frame, channel scaling, DSP impact, MCU feasibility matrix
PERFORMANCE-PROFILE.md DSP acceleration matrix, per-target CPU/memory budgets, CMSIS-DSP integration guide
PLATFORM-PICO.md RP2350 (Pico 2) hardware setup, build configuration, and XAP integration
PLATFORM-STM32.md STM32WB55/WBA55 configuration and CMSIS-DSP setup
PLATFORM-ESP32.md ESP32-S3 WiFi, ESP-IDF integration, PIE SIMD configuration
PLATFORM-NRF.md nRF9160 LTE-M setup and CMSIS-DSP integration
FEASIBILITY-RP2350.md Detailed feasibility analysis: 4ch @96kHz XAP on RP2350, CPU/memory budget breakdown
SDK-GUIDE.md Embedded SDK architecture; codec section cross-references this document

XAP (Xylolabs Audio Protocol) is a proprietary technology of Xylolabs Inc.

Patent applications covering the XAP encoding algorithm, frame format, adaptive quantization scheme, and platform-specific DSP acceleration paths have been filed. All rights reserved.

Unauthorized use, reproduction, reverse engineering, distribution, or incorporation of XAP or any portion thereof into third-party products or services is strictly prohibited without a written license from Xylolabs Inc.

For licensing inquiries, contact: legal@xylolabs.com

Trademarks: "Xylolabs", "XAP", and "XMBP" are trademarks of Xylolabs Inc.


Xylolabs Inc. — XAP Specification — Revision 2026-07-08