AI Foundry lab · Troubleshooting · 25 September 2026

What is broken on aifoundry1

Fixed on 25 September 2026, 15:02 PDT. Both cards answer again; what was done, and what is still open: aifoundry1 is fixed. Update, 27 September: card 1 has since run the three-card version-3 check (25–26 September) and the gathers and scatters after it; card 0 overheats under load (98–102 °C in short smoke launches, 115–117 °C idle afterwards) and takes no sustained work.

The cards are refused by software, not by hardware. Both ET-SoC-1 cards in aifoundry1 are on the bus, bound to the ET driver, and have their device nodes. What fails is a version check: the et_soc1 kernel module that DKMS builds on aifoundry1 has an empty version string (/sys/module/et_soc1/version holds only a newline, where aifoundry2 and aifoundry3 read 0.20.0), so et-platform’s deviceLayer, inside every ET program, throws Error unable to evaluate compatibility! before a single command reaches a card. The cause is a one-character Makefile typo, $(ET_MODULE_VERSION=), in the stale driver copy DKMS builds from (section 3), so a reboot will not help. What we cannot know until the fix is in is which firmware the cards run and whether they answer.

Broken
The et_soc1 module’s version string is empty, so deviceLayer refuses both cards.
Why
DKMS builds from a stale copy of 09531e5c1, whose Makefile has the typo. The fixed source (353f20e) is on the machine but DKMS was never pointed at it.
Fix
Root, about 5 minutes, with nobody using the cards: replace the DKMS source, rebuild for the three kernels in /boot, reload et_soc1 (section 1).
Sure?
High for the cause. The cards’ firmware and state are unverified until the fix.
Also found
Card et0’s PCIe link logs about one corrected error per second. It does not block the cards, but it is what fills /var/log/kern.log (section 4).
And
The root file system has about 164 MB left, and the 24 September kernel update stopped halfway: 7.0.0-34 has no initramfs and GRUB was not updated (section 5).
Correction to our earlier note

Correction to our earlier note (“a kernel-module srcversion mismatch with libDM.so”, in docs/findings/14-card-behaviour.md): it was wrong on both counts. Nothing in the ET stack checks srcversion, and the check is not in libDM.so. The evidence was already in our 22 September output, where the version line on aifoundry1 printed empty.

1. The fix (needs root on aifoundry1)

This rebuilds only the et_soc1 kernel module. Leave /opt/et alone: libDM.so and dev_mngt_service are byte-identical to the working machines. Run everything in one root shell (sudo -i), because later steps use the backup directory $B set in step 2. Keep steps 3 to 5 together: between them no et-soc1 module is on disk, so do not reboot in the middle.

The seven steps, with their commands
  1. Check that nobody is using the cards, and pause CI.

    Another user has been logged in since 18 September but holds no card: the module’s reference count is 0. The GitHub Actions runner for aifoundry1-et-soc1 runs as root; once the cards open, its jobs will be able to use them.

    who
    fuser -v /dev/et0_* /dev/et1_*        # expect no processes
    cat /sys/module/et_soc1/refcnt        # expect 0
    pgrep -af Runner.Worker               # expect nothing: no CI job running
    df -h /                               # 164 MB left on 25 Sep; steps 3-4 need ~20 MB
    systemctl stop actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1.service
  2. Back up the current module files and state.

    The old source directory is moved into the same backup directory in step 3.

    B=/root/et-soc1-fix-$(date +%Y%m%d-%H%M); mkdir -p $B
    dkms status > $B/dkms-status.before
    modinfo et_soc1 > $B/modinfo.before
    for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do
      cp -a /lib/modules/$k/updates/dkms/et-soc1.ko.zst $B/et-soc1.ko.zst.$k
    done
  3. Replace the DKMS source with the fixed driver that is already on the machine.

    The running module is not touched. dkms remove must come before the move, because DKMS needs the old dkms.conf to remove. git archive gives a clean tree without the January 2026 build leftovers in the checkout. root owns /usr/src/et-platform, so git works there for root without extra settings.

    git -C /usr/src/et-platform rev-parse HEAD      # expect 353f20e982f4...
    dkms remove et-soc1/0.20.0 --all
    mv /usr/src/et-soc1-0.20.0 $B/et-soc1-0.20.0.09531e5c1
    mkdir /usr/src/et-soc1-0.20.0
    git -C /usr/src/et-platform archive 353f20e et-driver |
      tar -x --strip-components=1 -C /usr/src/et-soc1-0.20.0
    sed -n 4,5p /usr/src/et-soc1-0.20.0/Makefile   # expect $(ET_MODULE_VERSION), no '=' inside
    cat /usr/src/et-soc1-0.20.0/VERSION            # expect 0.20.0
  4. Register it and build for every kernel in /boot.

    /boot holds 7.0.0-30, 7.0.0-31 (running) and 7.0.0-34 (installed on 24 September, but its install stopped halfway: section 5). Its headers are complete, so the build works. The --all above also drops the builds for 17 old 6.8.0 kernels that are no longer installed (no kernel image; headers remain only for 6.8.0-139); they need no rebuild. Secure Boot is off (mokutil --sb-state), so no key enrolment is needed.

    dkms add et-soc1/0.20.0
    for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do
      dkms install et-soc1/0.20.0 -k $k
    done
  5. Check the new files before touching the running module.

    Every line must read [0.20.0] 47D26A305A0428B29FB7FC4, the same as aifoundry2 and aifoundry3. If not, stop and roll back (below).

    for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do
      echo "$k [$(modinfo -k $k -F version et_soc1)] $(modinfo -k $k -F srcversion et_soc1)"
    done
  6. Reload the module.

    modprobe -r refuses while any device node is open. The driver’s remove path only releases host-side resources (interrupts, memory regions, bus mastering); it does not reset the cards, so their firmware keeps running and the new module re-reads its ready status. Load it by the name et_soc1: an older esperanto driver on disk drives the same cards.

    modprobe -r et_soc1 && modprobe et_soc1 && udevadm settle
    ls -l /dev/et*                                  # four nodes, crw-rw-rw-
  7. Load it at boot the same way as aifoundry2, and restart CI.

    Today nothing visible loads et_soc1 at boot; something does so 3 min 42 s after start-up (see section 6). Please also check the runner’s workflow for steps that run make dkms, insmod or modprobe from another checkout, since such a step could undo this fix.

    echo et_soc1 > /etc/modules-load.d/et_soc1.conf
    systemctl start actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1.service
Rollback

Rollback

This puts back the same broken module as today. If you are in a new shell, first set B to the backup directory from step 2.

modprobe -r et_soc1
rm -f /etc/modules-load.d/et_soc1.conf
dkms remove et-soc1/0.20.0 --all
rm -rf /usr/src/et-soc1-0.20.0
cp -a $B/et-soc1-0.20.0.09531e5c1 /usr/src/et-soc1-0.20.0
dkms add et-soc1/0.20.0
for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do dkms install et-soc1/0.20.0 -k $k; done
modprobe et_soc1
Minimal alternative

Minimal alternative

If you would rather keep the old directory, do steps 1 and 2, then replace only its Makefile with the one from the upstream fix: between 09531e5c1 and 78ed9b0d6 the Makefile differs in exactly the two version lines. Then run steps 5 to 7, but in step 5 expect [0.20.0] 1383B256EB24A0A53F04CC7: the version becomes 0.20.0 (we built this variant to check), while srcversion stays that of the old source, which is harmless but differs from aifoundry2 and aifoundry3.

cp /usr/src/et-soc1-0.20.0/Makefile $B/Makefile.09531e5c1
git -C /usr/src/et-platform show 78ed9b0d6:et-driver/Makefile > /usr/src/et-soc1-0.20.0/Makefile
for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do
  dkms build et-soc1/0.20.0 -k $k --force && dkms install et-soc1/0.20.0 -k $k --force
done

Do not load the older esperanto.ko as a shortcut: its version string would pass, but it is an older driver that aifoundry2 and aifoundry3 do not run. Optional cleanup, as a separate change: dkms remove esperanto/0.20.0 --all and remove /etc/udev/rules.d/50-esperanto.rules. A user-level stopgap without root is possible (an LD_PRELOAD shim that makes deviceLayer read 0.20.0), but we have not tried it on a card, and the root fix is quicker.

2. How to check that it worked

The full checklist
  1. The version deviceLayer reads is back.

    Expect 0.20.0 three times, then 47D26A305A0428B29FB7FC4. The two PCI paths are the exact files deviceLayer reads.

    cat /sys/module/et_soc1/version
    cat /sys/bus/pci/devices/0000:01:00.0/driver/module/version
    cat /sys/bus/pci/devices/0000:02:00.0/driver/module/version
    cat /sys/module/et_soc1/srcversion
  2. The other kernels are fixed too. modinfo -k 7.0.0-34-generic -F version et_soc1 prints 0.20.0, and so does the same command for 7.0.0-30-generic.
  3. A card answers. With nobody else using the cards (any user can run this):
    timeout 10 /opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n 0 -u 5000
    timeout 10 /opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n 1 -u 5000

    Pass: Service request succeeded and no evaluate compatibility. For comparison, aifoundry2’s card reported today: firmware release 1.3.1, BL1 and BL2 0.20.0, PMIC 1.5.0, master, worker and machine minion 0.23.0. dev_mngt_service opens both cards’ management nodes whatever -n says, so a fault on either card fails both runs.

  4. Commands reached both cards. cat /sys/bus/pci/devices/0000:0{1,2}:00.0/mgmt_vq_stats/msg_count: the SQ0 count should be at least 1 on each card. Today it is 0 on both.
  5. After the next reboot (whichever kernel it boots; see section 5): lsmod | grep et_soc1 shows the module, loaded by modules-load.d, and checks 1 and 3 pass again.

What would not prove anything, or would be a different problem:

  • ls -l /dev/et* showing 0666 and lspci -k showing Kernel driver in use: ET are already true today.
  • Do not use DM_CMD_GET_FIRMWARE_BOOT_STATUS as a test. It fails with Received incorrect rsp status: -16006 on the working aifoundry3 too.
  • Error device /dev/etN_mgmt is in bad state, state: N would mean the driver fix worked and the card or firmware state is the next problem (the device-state check right after the version check, DevicePcie.cpp line 298).
  • Incompatible device-api version from a runtime program would mean the card firmware does not match the runtime, which needs device-api 2.x (aifoundry2 reports 2.4.0).

3. Evidence

The error, verbatim

The error, verbatim

From our attempts on aifoundry1 on 22 September 2026, with the stack traces and log prefixes shortened. Our telemetry tool, which links deviceLayer:

Opening device 0
FAIL: Exception message:Error unable to evaluate compatibility!
StackTrace:
	stack dump [1]  dbg::StackException::StackException(std::__cxx11::basic_string<...> const&) + 0x46
	...

The lab’s own tool, /opt/et/bin/dev_mngt_service -m DM_CMD_GET_FIRMWARE_BOOT_STATUS -n 0, at 12:11:51 PDT:

command: DM_CMD_GET_FIRMWARE_BOOT_STATUS code 16
terminate called after throwing an instance of 'dev::Exception'
  what():  Exception message:Error unable to evaluate compatibility!
StackTrace:
2026/09/22 12:11:51 157256
***** FATAL SIGNAL RECEIVED *******

In the same command, cat /sys/module/et_soc1/version printed an empty line on aifoundry1 and 0.20.0 on aifoundry3. The same dev_mngt_service binary on aifoundry3 got past the check and received a reply from its card (status -16006, which this command also gets on working cards). Our raw-ioctl tool, which does not use deviceLayer, opened both aifoundry1 cards the same day and read their configuration (65 W TDP, 600 MHz boot clock, all 32 shires), identical to aifoundry3’s.

Where the check is

Where the check is

et-platform devicelayer/src/DevicePcie.cpp, function openWhenReady(), shortened below with our comments. This function is the same at 353f20e, from which dev_mngt_service and et-powertop were built on all three machines; in the May 2026 local rebuild of libdeviceLayer.a on aifoundry1 (commit acc7ed25d in /home/rehan/et-platform, which changes other parts of this file); and at 836a4ab60, the newest upstream commit in our clone.

 36  constexpr auto kMinReqDriverVersion = "0.15.0";
     ...
254  int openWhenReady(const std::string path, std::chrono::seconds timeout) {
       // open(path, O_RDWR | O_NONBLOCK), then ioctl ETSOC1_IOCTL_GET_PCIBUS_DEVICE_NAME -> "0000:01:00.0"
270    auto curVersion = getDeviceAttributeByName(devName, "driver/module/version");
         // reads /sys/bus/pci/devices/0000:01:00.0/driver/module/version  ->  "\n" on aifoundry1
274    if (std::regex rgx("(0|[1-9][0-9]*)\\.(0|[1-9][0-9]*)\\.(0|[1-9][0-9]*)"); std::regex_search(curVersion, ...)) {
         // major must be 0 and the version at least 0.15.0
       } else {
         // Driver does not follow semantic versioning
294      throw Exception("Error unable to evaluate compatibility!");
       }
298    // device-state check ("... is in bad state ..."): never reached on aifoundry1
  • deviceLayer is linked statically into dev_mngt_service, et-powertop and every program built on libdeviceLayer.a. The strings evaluate compatibility and driver/module/version appear in those binaries on both aifoundry1 and aifoundry2, and in neither libDM.so nor libetrt.so.
  • git grep srcversion over the whole et-platform tree finds nothing.
  • The empty-version branch has been there since af48d9e43 (October 2022). The upstream fix 78ed9b0d6 (26 October 2025, “[et-driver] Fix version on dkms build”) says: “When version is empty, devicelayer fails to open the device as cannot check the compatibility of the driver.”
The Makefile typo

The Makefile typo

aifoundry1, /usr/src/et-soc1-0.20.0/Makefile lines 4-5 (= et-platform 09531e5c1):
  ET_MODULE_VERSION=$(shell cat VERSION)
  CFLAGS_MODULE+=-DET_MODULE_VERSION=\"$(ET_MODULE_VERSION=)\"     ->  -DET_MODULE_VERSION=\"\"

aifoundry2 and aifoundry3 (78ed9b0d6 and later):
  ET_MODULE_VERSION := $(or $(shell cat VERSION 2>/dev/null), $(shell cat $(src)/VERSION 2>/dev/null), unknown)
  CFLAGS_MODULE+=-DET_MODULE_VERSION=\"$(ET_MODULE_VERSION)\"      ->  -DET_MODULE_VERSION=\"0.20.0\"

$(ET_MODULE_VERSION=) names a make variable called ET_MODULE_VERSION=, which does not exist, so it expands to nothing. The driver then declares MODULE_VERSION(""), so modinfo shows an empty version: and the sysfs file holds one byte, a newline.

Side by side

Observed on 25 September 2026, 13:01–14:05 PDT, except where marked. Shaded cells are where aifoundry1 differs in a way that matters.

aifoundry1aifoundry2aifoundry3
Kernel running / newest installed7.0.0-31 / 7.0.0-34, install stopped halfway: no initramfs, GRUB not updated7.0.0-31 / 7.0.0-34, completesame as aifoundry2
Space left on /about 164 MB (99% used)229 GiB342 GiB
ET cards on PCIe2: 01:00.0 (et0), 02:00.0 (et1)1: 02:00.01: 02:00.0
Driver bound, linkET on both, Gen4 16 GT/s x8samesame
Device nodeset0_mgmt, et0_ops, et1_mgmt, et1_ops; 0666; created 18 Sep 15:47:31–32et0_mgmt, et0_ops; 0666same as aifoundry2
/sys/module/et_soc1/versionempty (1 byte, a newline)0.20.00.20.0
srcversion (checked by nothing)1383B256EB24A0A53F04CC747D26A305A0428B29FB7FC447D26A305A0428B29FB7FC4
DKMS source /usr/src/et-soc1-0.20.0exact copy of 09531e5c1 (21 Oct 2025); files dated 24 Oct 2025= 353f20e; dated 12 Mar 2026= 353f20e; dated 25 Mar 2026
Makefile version line$(ET_MODULE_VERSION=) (typo)fixedfixed
Module built for 7.0.0-34version empty0.20.00.20.0
Compiled driver codesame as aifoundry2’s: 48 of 51 ELF sections byte-identical (all code, data and relocations); only the module-info strings (version, srcversion), symbol offsets, build ID and the per-host signature differreferencesame source as aifoundry2
Fixed driver source on diskyes: /usr/src/et-platform at 353f20e, not used by DKMSin usein use
libDM.so, dev_mngt_servicesha256 28712bc7…, a3d4c722…identicalidentical
libdeviceLayer.a (holds the check)rebuilt 1 May 2026 from /home/rehan/et-platform (commit acc7ed25d, which adds an ET_DEVICES device filter); check present3 Jan 20263 Jan 2026
What loads et_soc1 at bootnot visible to a user; loaded at boot +3 min 42 s; no modules-load.d entry, no udev alias/etc/modules-load.d/et_soc1.conf (boot +7 s)etsoc1-demo.service runs modprobe (boot +15 s) inferred
Commands sent to the card this boot (mgmt SQ0)0 on both cards3.25 million1.63 million
Driver error counterset0: one PMIC correctable event; et1: none; no uncorrectable errorsMinionCe 23, SpCe 5SpCe 3
Corrected PCIe errors at the card’s root port, this bootet0 (00:01.0): about 695,000; et1 (00:01.1): 030
/var/log/kern.logabout 280 MB0.19 MB0.08 MB
Other ET driver on diskolder esperanto 0.20.0 DKMS package, built for 23 kernels, not loadednonenone
Others on the machineanother user, logged in since 18 Sep; root GitHub Actions runnerour experiment queue; another user; a root GitHub Actions runnerour sampler; the etsoc1-demo and etsoc1-chatbot services
Reproduced from source

Reproduced from source

We built the driver three ways on aifoundry2, in a scratch directory against the 7.0.0-31 headers, and installed nothing:

BuildversionsrcversionMatches
A: 09531e5c1 as it isempty1383B256…aifoundry1’s installed module
B: 09531e5c1 with the Makefile from 78ed9b0d60.20.01383B256…the minimal alternative in section 1
C: 353f20e0.20.047D26A30…aifoundry2’s and aifoundry3’s installed module

The code and data sections are identical in all three builds and in both installed modules. srcversion is a checksum of the driver’s .c and .h text, not of its Makefile. It differs only because commit 307e59c27 added a CentOS 9 condition, && (!defined(RHEL_MAJOR) || RHEL_MAJOR < 9), to five #if lines in et-soc1-pcie.c and et_pci_dev.h (and one in the loopback variant, which is not built). RHEL_MAJOR is not defined on Ubuntu, so the compiled code does not change. Nothing in the ET stack reads srcversion; DKMS uses it to compare builds.

How it got this way

This does not stop the cards from opening, and the module fix does not touch it. The root port above card et0 (00:01.0, the CPU’s x16 root port, leading to 01:00.0) had counted 695,018 corrected receiver errors (RxErr) and 5 bad DLLPs this boot at 13:34. They keep arriving at 1.0–1.4 per second while the card is idle, at random intervals. There are no non-fatal or fatal errors, and the link still runs at 16 GT/s x8. The errors are on the card-to-CPU direction: et0’s own counters are 0. For comparison, et1’s port (00:01.1) has 0, aifoundry2’s card has 3 for its whole boot under heavy load, and aifoundry3’s has 0. That is a bit error rate of at least about 10−11, roughly ten times the PCIe target: a marginal link.

Full detail: log growth, whether it varies between boots, the risk, and what to check
  • It is what fills the logs. We sampled once a second for 97 seconds. kern.log never grew in a second without a new error. In 60 of the 70 seconds with errors it grew by exactly 527 bytes per error, the size of one four-line AER report; in the rest, the kernel’s rate limit (10 reports per 5 s) dropped some reports. syslog grows the same way. That is about 50 MB a day in each file, which is why kern.log is about 280 MB and the journal now reaches back only to 22 September. The root file system has about 164 MB left (section 5).
  • It varies between boots inferred. From the compressed sizes of older logs, the stream was absent in the 30 August to 18 September boot and may have been present in a late-August boot. So one clean boot after a reseat is not proof of a fix.
  • The risk inferred from the code. If the link ever produces a fatal error, the ET driver gives et0 up: it answers the kernel’s error recovery with DISCONNECT and tears down et0’s queues, so et0 stays unusable until it is probed again, in practice at a reboot. dev_mngt_service and et-powertop open every card’s management node at start-up, starting with et0, so they would then fail for et1 too. Programs built on aifoundry1 against its locally rebuilt libdeviceLayer.a can be limited to one card with ET_DEVICES=1, a local addition that is not in upstream et-platform.

What we suggest (root):

# limit the log to 10 reports an hour; the error counters keep counting; reverts at reboot
echo 3600000 > /sys/bus/pci/devices/0000:00:01.0/aer/correctable_ratelimit_interval_ms
# record the link detail only root can read (ASPM, link status, error masks, lane errors)
lspci -vvv -s 00:01.0; lspci -vvv -s 01:00.0
# watch the rate; the TOTAL_ERR_COR line should keep rising while kern.log stays quiet
cat /sys/bus/pci/devices/0000:00:01.0/aer_dev_correctable

At a maintenance window: after the module fix, note each card’s serial number (/dev/etN follows probe order). Then power off, reseat et0 and check its power connector. If the errors persist, swap et0 and et1: errors that follow the card mean the card, errors that stay on 00:01.0 mean the slot or board. Forcing that slot to Gen3 in the BIOS is a fallback. Judge the result by the hourly rate in aer_dev_correctable over at least two boots, not by the size of kern.log.

5. Also: a nearly full disk and a half-finished kernel update

Neither of these blocks the cards today, but both matter for the fix and for the next reboot.

The disk, the interrupted kernel update, and what to check
  • The root file system is nearly full. / has about 164 MB left (99% used, 13:57). The pool is 96% allocated, mostly by /home (421 GB) and /root (27 GB). /var/log takes 268 MB on disk (170 MB of it the journal), and the AER reports from section 4 keep adding to it. aifoundry2 and aifoundry3 have hundreds of GB free.
  • The 24 September kernel update stopped partway. dpkg lists linux-image-7.0.0-34-generic as half-configured (iF). Its last dpkg.log entry, at 06:08:51, is the trigger that builds the initramfs and updates GRUB. After that:
    • /boot has vmlinuz-7.0.0-34-generic but no initrd.img-7.0.0-34-generic, and the /boot/initrd.img link points at the missing file;
    • /boot/grub/grub.cfg was last written on 5 September;
    • /var/lib/dpkg/updates still holds entries from the interrupted run, and apt’s history.log has no end line for it.
    On aifoundry2 and aifoundry3 the same update finished, with the initramfs and grub.cfg written at 06:08 and 06:30.
  • What it means inferred. The next boot most likely still starts 7.0.0-31 from the September grub.cfg. Booting 7.0.0-34 without an initramfs would not mount the ZFS root. A lack of space is a likely cause of the stop, but the installer’s output (/var/log/apt/term.log) is readable only by root and adm, so we could not confirm it. The module fix in section 1 already builds et_soc1 for 7.0.0-34, so it does not depend on this.

What we suggest (root), after freeing some space:

df -h /                                           # make room first
dpkg --configure -a                               # finishes the 7.0.0-34 install: initramfs and GRUB
ls -l /boot/initrd.img-7.0.0-34-generic           # should now exist
dpkg -l linux-image-7.0.0-34-generic | tail -1    # should start with ii

6. What we did not check or could not see without root

Caveats in full
  • The cards’ firmware and state. No command has reached either card this boot, so their firmware versions, their device state and the runtime’s device-API check are unknown until the fix. The device nodes exist, and the driver creates them only after the firmware reports ready, so both cards were up when the module loaded inferred from the driver code. The firmware images in aifoundry1’s /opt/et are a local build (acc7ed25d-dirty, 10 May 2026), and /opt/etsw holds another set (d08fcdcb); aifoundry2 and aifoundry3 have 353f20e images. Images on disk need not match what is flashed: aifoundry2’s disk has BL 0.22.0 and minion 0.25.0 images, but its card reports BL 0.20.0 and minion 0.23.0. Which firmware aifoundry1’s cards run is unknown.
  • The kernel log. dmesg and the journal are closed to us, so we did not see the probe-time messages, the one PMIC correctable event on et0, or when the RxErr stream started. /var/log/dmesg (148 MB, readable by root and adm), written on 22 September by journalctl --boot 0 --dmesg, probably holds this boot’s kernel log up to that day, including the 18 September probe inferred from its size.
  • What loads et_soc1 at boot +3 min 42 s. It is not modules-load.d or udev. The candidates are the root GitHub Actions runner, a root cron job or a person inferred; /root is not readable, so we could not read the runner’s workflow either.
  • Whether the hand-built fixed modules from January and July 2026 were ever loaded.
  • The et0 link detail (ASPM, per-lane errors, error masks): PCIe config space beyond 64 bytes is root-only. Whether the card or the slot is at fault needs a reseat or a swap.
  • Runtime parity after the fix. aifoundry1’s libdeviceLayer.a (1 May 2026) and libetrt.so (10 May 2026) were rebuilt from /home/rehan/et-platform and differ from aifoundry2’s. They are not part of this fault, but runtime programs on aifoundry1 will not run the same library binaries as on aifoundry2.

7. How the facts were gathered

Method, in full
  • When. Friday 25 September 2026, 13:01–14:05 PDT. All three machines have been up since 18 September, 15:43.
  • Read-only. As user yaroslavvb over SSH, with no sudo: sysfs, /proc, modinfo, dkms status, lspci as a user, file listings and checksums, git in our clone of et-platform, and read-only git queries in the two et-platform checkouts on aifoundry1. Nothing was installed, and no module was loaded or unloaded. On aifoundry1 the only writes were two temporary decompressed copies of the module in our home directory, deleted at once.
  • No card was opened. who on aifoundry1 showed another user logged in at every check, from 13:01 to 14:05, so we opened no /dev/et* node there. On aifoundry2 and aifoundry3 our own experiment queues hold the cards, so we read only sysfs and files.
  • The error text is from our attempts on 22 September 2026. They failed before any command reached a card; the management queue counters are still 0.
  • The rebuilds ran on aifoundry2 in a scratch directory (make -C /lib/modules/7.0.0-31-generic/build M=<dir> modules). The installed modules of aifoundry1 and aifoundry2 were compared section by section.
  • Checked twice. Each conclusion was checked by a second, independent pass along three lines: the driver and its interface, the hardware, and how the driver was deployed. All three confirmed the cause. The second passes corrected details, not the conclusion: the verification commands, the log-flood link and the start date of the error stream. A last pass re-ran the cheap checks against this page (13:54–14:05) and corrected the kernel-headers claim for old kernels, the #if count, the state of the 7.0.0-34 install and the ET_DEVICES filter in aifoundry1’s rebuilt deviceLayer.