AI Foundry lab · Troubleshooting · 25 September 2026
What is broken on aifoundry1
Fixed on 25 September 2026, 15:02 PDT. Both cards answer again; what was done, and what is still open: aifoundry1 is fixed. Update, 27 September: card 1 has since run the three-card version-3 check (25–26 September) and the gathers and scatters after it; card 0 overheats under load (98–102 °C in short smoke launches, 115–117 °C idle afterwards) and takes no sustained work.
The cards are refused by software, not by hardware. Both ET-SoC-1 cards in aifoundry1 are on the bus, bound to the ET driver, and have their device nodes. What fails is a version check: the et_soc1 kernel module that DKMS builds on aifoundry1 has an empty version string (/sys/module/et_soc1/version holds only a newline, where aifoundry2 and aifoundry3 read 0.20.0), so et-platform’s deviceLayer, inside every ET program, throws Error unable to evaluate compatibility! before a single command reaches a card. The cause is a one-character Makefile typo, $(ET_MODULE_VERSION=), in the stale driver copy DKMS builds from (section 3), so a reboot will not help. What we cannot know until the fix is in is which firmware the cards run and whether they answer.
- Broken
- The
et_soc1module’s version string is empty, so deviceLayer refuses both cards. - Why
- DKMS builds from a stale copy of
09531e5c1, whose Makefile has the typo. The fixed source (353f20e) is on the machine but DKMS was never pointed at it. - Fix
- Root, about 5 minutes, with nobody using the cards: replace the DKMS source, rebuild for the three kernels in
/boot, reloadet_soc1(section 1). - Sure?
- High for the cause. The cards’ firmware and state are unverified until the fix.
- Also found
- Card et0’s PCIe link logs about one corrected error per second. It does not block the cards, but it is what fills
/var/log/kern.log(section 4). - And
- The root file system has about 164 MB left, and the 24 September kernel update stopped halfway:
7.0.0-34has no initramfs and GRUB was not updated (section 5).
Correction to our earlier note
Correction to our earlier note (“a kernel-module srcversion mismatch with libDM.so”, in docs/findings/14-card-behaviour.md): it was wrong on both counts. Nothing in the ET stack checks srcversion, and the check is not in libDM.so. The evidence was already in our 22 September output, where the version line on aifoundry1 printed empty.
1. The fix (needs root on aifoundry1)
This rebuilds only the et_soc1 kernel module. Leave /opt/et alone: libDM.so and dev_mngt_service are byte-identical to the working machines. Run everything in one root shell (sudo -i), because later steps use the backup directory $B set in step 2. Keep steps 3 to 5 together: between them no et-soc1 module is on disk, so do not reboot in the middle.
The seven steps, with their commands
- Check that nobody is using the cards, and pause CI.
Another user has been logged in since 18 September but holds no card: the module’s reference count is 0. The GitHub Actions runner for
aifoundry1-et-soc1runs as root; once the cards open, its jobs will be able to use them.who fuser -v /dev/et0_* /dev/et1_* # expect no processes cat /sys/module/et_soc1/refcnt # expect 0 pgrep -af Runner.Worker # expect nothing: no CI job running df -h / # 164 MB left on 25 Sep; steps 3-4 need ~20 MB systemctl stop actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1.service - Back up the current module files and state.
The old source directory is moved into the same backup directory in step 3.
B=/root/et-soc1-fix-$(date +%Y%m%d-%H%M); mkdir -p $B dkms status > $B/dkms-status.before modinfo et_soc1 > $B/modinfo.before for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do cp -a /lib/modules/$k/updates/dkms/et-soc1.ko.zst $B/et-soc1.ko.zst.$k done - Replace the DKMS source with the fixed driver that is already on the machine.
The running module is not touched.
dkms removemust come before the move, because DKMS needs the olddkms.confto remove.git archivegives a clean tree without the January 2026 build leftovers in the checkout. root owns/usr/src/et-platform, so git works there for root without extra settings.git -C /usr/src/et-platform rev-parse HEAD # expect 353f20e982f4... dkms remove et-soc1/0.20.0 --all mv /usr/src/et-soc1-0.20.0 $B/et-soc1-0.20.0.09531e5c1 mkdir /usr/src/et-soc1-0.20.0 git -C /usr/src/et-platform archive 353f20e et-driver | tar -x --strip-components=1 -C /usr/src/et-soc1-0.20.0 sed -n 4,5p /usr/src/et-soc1-0.20.0/Makefile # expect $(ET_MODULE_VERSION), no '=' inside cat /usr/src/et-soc1-0.20.0/VERSION # expect 0.20.0 - Register it and build for every kernel in
/boot./bootholds 7.0.0-30, 7.0.0-31 (running) and 7.0.0-34 (installed on 24 September, but its install stopped halfway: section 5). Its headers are complete, so the build works. The--allabove also drops the builds for 17 old 6.8.0 kernels that are no longer installed (no kernel image; headers remain only for 6.8.0-139); they need no rebuild. Secure Boot is off (mokutil --sb-state), so no key enrolment is needed.dkms add et-soc1/0.20.0 for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do dkms install et-soc1/0.20.0 -k $k done - Check the new files before touching the running module.
Every line must read
[0.20.0] 47D26A305A0428B29FB7FC4, the same as aifoundry2 and aifoundry3. If not, stop and roll back (below).for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do echo "$k [$(modinfo -k $k -F version et_soc1)] $(modinfo -k $k -F srcversion et_soc1)" done - Reload the module.
modprobe -rrefuses while any device node is open. The driver’s remove path only releases host-side resources (interrupts, memory regions, bus mastering); it does not reset the cards, so their firmware keeps running and the new module re-reads its ready status. Load it by the nameet_soc1: an olderesperantodriver on disk drives the same cards.modprobe -r et_soc1 && modprobe et_soc1 && udevadm settle ls -l /dev/et* # four nodes, crw-rw-rw- - Load it at boot the same way as aifoundry2, and restart CI.
Today nothing visible loads
et_soc1at boot; something does so 3 min 42 s after start-up (see section 6). Please also check the runner’s workflow for steps that runmake dkms,insmodormodprobefrom another checkout, since such a step could undo this fix.echo et_soc1 > /etc/modules-load.d/et_soc1.conf systemctl start actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1.service
Rollback
Rollback
This puts back the same broken module as today. If you are in a new shell, first set B to the backup directory from step 2.
modprobe -r et_soc1
rm -f /etc/modules-load.d/et_soc1.conf
dkms remove et-soc1/0.20.0 --all
rm -rf /usr/src/et-soc1-0.20.0
cp -a $B/et-soc1-0.20.0.09531e5c1 /usr/src/et-soc1-0.20.0
dkms add et-soc1/0.20.0
for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do dkms install et-soc1/0.20.0 -k $k; done
modprobe et_soc1Minimal alternative
Minimal alternative
If you would rather keep the old directory, do steps 1 and 2, then replace only its Makefile with the one from the upstream fix: between 09531e5c1 and 78ed9b0d6 the Makefile differs in exactly the two version lines. Then run steps 5 to 7, but in step 5 expect [0.20.0] 1383B256EB24A0A53F04CC7: the version becomes 0.20.0 (we built this variant to check), while srcversion stays that of the old source, which is harmless but differs from aifoundry2 and aifoundry3.
cp /usr/src/et-soc1-0.20.0/Makefile $B/Makefile.09531e5c1
git -C /usr/src/et-platform show 78ed9b0d6:et-driver/Makefile > /usr/src/et-soc1-0.20.0/Makefile
for k in 7.0.0-30-generic 7.0.0-31-generic 7.0.0-34-generic; do
dkms build et-soc1/0.20.0 -k $k --force && dkms install et-soc1/0.20.0 -k $k --force
doneDo not load the older esperanto.ko as a shortcut: its version string would pass, but it is an older driver that aifoundry2 and aifoundry3 do not run. Optional cleanup, as a separate change: dkms remove esperanto/0.20.0 --all and remove /etc/udev/rules.d/50-esperanto.rules. A user-level stopgap without root is possible (an LD_PRELOAD shim that makes deviceLayer read 0.20.0), but we have not tried it on a card, and the root fix is quicker.
2. How to check that it worked
The full checklist
- The version deviceLayer reads is back.
Expect
0.20.0three times, then47D26A305A0428B29FB7FC4. The two PCI paths are the exact files deviceLayer reads.cat /sys/module/et_soc1/version cat /sys/bus/pci/devices/0000:01:00.0/driver/module/version cat /sys/bus/pci/devices/0000:02:00.0/driver/module/version cat /sys/module/et_soc1/srcversion - The other kernels are fixed too.
modinfo -k 7.0.0-34-generic -F version et_soc1prints0.20.0, and so does the same command for7.0.0-30-generic. - A card answers. With nobody else using the cards (any user can run this):
timeout 10 /opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n 0 -u 5000 timeout 10 /opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n 1 -u 5000Pass:
Service request succeededand noevaluate compatibility. For comparison, aifoundry2’s card reported today: firmware release 1.3.1, BL1 and BL2 0.20.0, PMIC 1.5.0, master, worker and machine minion 0.23.0.dev_mngt_serviceopens both cards’ management nodes whatever-nsays, so a fault on either card fails both runs. - Commands reached both cards.
cat /sys/bus/pci/devices/0000:0{1,2}:00.0/mgmt_vq_stats/msg_count: the SQ0 count should be at least 1 on each card. Today it is 0 on both. - After the next reboot (whichever kernel it boots; see section 5):
lsmod | grep et_soc1shows the module, loaded bymodules-load.d, and checks 1 and 3 pass again.
What would not prove anything, or would be a different problem:
ls -l /dev/et*showing 0666 andlspci -kshowingKernel driver in use: ETare already true today.- Do not use
DM_CMD_GET_FIRMWARE_BOOT_STATUSas a test. It fails withReceived incorrect rsp status: -16006on the working aifoundry3 too. Error device /dev/etN_mgmt is in bad state, state: Nwould mean the driver fix worked and the card or firmware state is the next problem (the device-state check right after the version check,DevicePcie.cppline 298).Incompatible device-api versionfrom a runtime program would mean the card firmware does not match the runtime, which needs device-api 2.x (aifoundry2 reports 2.4.0).
3. Evidence
The error, verbatim
The error, verbatim
From our attempts on aifoundry1 on 22 September 2026, with the stack traces and log prefixes shortened. Our telemetry tool, which links deviceLayer:
Opening device 0
FAIL: Exception message:Error unable to evaluate compatibility!
StackTrace:
stack dump [1] dbg::StackException::StackException(std::__cxx11::basic_string<...> const&) + 0x46
...
The lab’s own tool, /opt/et/bin/dev_mngt_service -m DM_CMD_GET_FIRMWARE_BOOT_STATUS -n 0, at 12:11:51 PDT:
command: DM_CMD_GET_FIRMWARE_BOOT_STATUS code 16
terminate called after throwing an instance of 'dev::Exception'
what(): Exception message:Error unable to evaluate compatibility!
StackTrace:
2026/09/22 12:11:51 157256
***** FATAL SIGNAL RECEIVED *******
In the same command, cat /sys/module/et_soc1/version printed an empty line on aifoundry1 and 0.20.0 on aifoundry3. The same dev_mngt_service binary on aifoundry3 got past the check and received a reply from its card (status -16006, which this command also gets on working cards). Our raw-ioctl tool, which does not use deviceLayer, opened both aifoundry1 cards the same day and read their configuration (65 W TDP, 600 MHz boot clock, all 32 shires), identical to aifoundry3’s.
Where the check is
Where the check is
et-platform devicelayer/src/DevicePcie.cpp, function openWhenReady(), shortened below with our comments. This function is the same at 353f20e, from which dev_mngt_service and et-powertop were built on all three machines; in the May 2026 local rebuild of libdeviceLayer.a on aifoundry1 (commit acc7ed25d in /home/rehan/et-platform, which changes other parts of this file); and at 836a4ab60, the newest upstream commit in our clone.
36 constexpr auto kMinReqDriverVersion = "0.15.0";
...
254 int openWhenReady(const std::string path, std::chrono::seconds timeout) {
// open(path, O_RDWR | O_NONBLOCK), then ioctl ETSOC1_IOCTL_GET_PCIBUS_DEVICE_NAME -> "0000:01:00.0"
270 auto curVersion = getDeviceAttributeByName(devName, "driver/module/version");
// reads /sys/bus/pci/devices/0000:01:00.0/driver/module/version -> "\n" on aifoundry1
274 if (std::regex rgx("(0|[1-9][0-9]*)\\.(0|[1-9][0-9]*)\\.(0|[1-9][0-9]*)"); std::regex_search(curVersion, ...)) {
// major must be 0 and the version at least 0.15.0
} else {
// Driver does not follow semantic versioning
294 throw Exception("Error unable to evaluate compatibility!");
}
298 // device-state check ("... is in bad state ..."): never reached on aifoundry1
- deviceLayer is linked statically into
dev_mngt_service,et-powertopand every program built onlibdeviceLayer.a. The stringsevaluate compatibilityanddriver/module/versionappear in those binaries on both aifoundry1 and aifoundry2, and in neitherlibDM.sonorlibetrt.so. git grep srcversionover the whole et-platform tree finds nothing.- The empty-version branch has been there since
af48d9e43(October 2022). The upstream fix78ed9b0d6(26 October 2025, “[et-driver] Fix version on dkms build”) says: “When version is empty, devicelayer fails to open the device as cannot check the compatibility of the driver.”
The Makefile typo
The Makefile typo
aifoundry1, /usr/src/et-soc1-0.20.0/Makefile lines 4-5 (= et-platform 09531e5c1):
ET_MODULE_VERSION=$(shell cat VERSION)
CFLAGS_MODULE+=-DET_MODULE_VERSION=\"$(ET_MODULE_VERSION=)\" -> -DET_MODULE_VERSION=\"\"
aifoundry2 and aifoundry3 (78ed9b0d6 and later):
ET_MODULE_VERSION := $(or $(shell cat VERSION 2>/dev/null), $(shell cat $(src)/VERSION 2>/dev/null), unknown)
CFLAGS_MODULE+=-DET_MODULE_VERSION=\"$(ET_MODULE_VERSION)\" -> -DET_MODULE_VERSION=\"0.20.0\"
$(ET_MODULE_VERSION=) names a make variable called ET_MODULE_VERSION=, which does not exist, so it expands to nothing. The driver then declares MODULE_VERSION(""), so modinfo shows an empty version: and the sysfs file holds one byte, a newline.
Side by side
Observed on 25 September 2026, 13:01–14:05 PDT, except where marked. Shaded cells are where aifoundry1 differs in a way that matters.
| aifoundry1 | aifoundry2 | aifoundry3 | |
|---|---|---|---|
| Kernel running / newest installed | 7.0.0-31 / 7.0.0-34, install stopped halfway: no initramfs, GRUB not updated | 7.0.0-31 / 7.0.0-34, complete | same as aifoundry2 |
Space left on / | about 164 MB (99% used) | 229 GiB | 342 GiB |
| ET cards on PCIe | 2: 01:00.0 (et0), 02:00.0 (et1) | 1: 02:00.0 | 1: 02:00.0 |
| Driver bound, link | ET on both, Gen4 16 GT/s x8 | same | same |
| Device nodes | et0_mgmt, et0_ops, et1_mgmt, et1_ops; 0666; created 18 Sep 15:47:31–32 | et0_mgmt, et0_ops; 0666 | same as aifoundry2 |
/sys/module/et_soc1/version | empty (1 byte, a newline) | 0.20.0 | 0.20.0 |
srcversion (checked by nothing) | 1383B256EB24A0A53F04CC7 | 47D26A305A0428B29FB7FC4 | 47D26A305A0428B29FB7FC4 |
DKMS source /usr/src/et-soc1-0.20.0 | exact copy of 09531e5c1 (21 Oct 2025); files dated 24 Oct 2025 | = 353f20e; dated 12 Mar 2026 | = 353f20e; dated 25 Mar 2026 |
| Makefile version line | $(ET_MODULE_VERSION=) (typo) | fixed | fixed |
| Module built for 7.0.0-34 | version empty | 0.20.0 | 0.20.0 |
| Compiled driver code | same as aifoundry2’s: 48 of 51 ELF sections byte-identical (all code, data and relocations); only the module-info strings (version, srcversion), symbol offsets, build ID and the per-host signature differ | reference | same source as aifoundry2 |
| Fixed driver source on disk | yes: /usr/src/et-platform at 353f20e, not used by DKMS | in use | in use |
libDM.so, dev_mngt_service | sha256 28712bc7…, a3d4c722… | identical | identical |
libdeviceLayer.a (holds the check) | rebuilt 1 May 2026 from /home/rehan/et-platform (commit acc7ed25d, which adds an ET_DEVICES device filter); check present | 3 Jan 2026 | 3 Jan 2026 |
What loads et_soc1 at boot | not visible to a user; loaded at boot +3 min 42 s; no modules-load.d entry, no udev alias | /etc/modules-load.d/et_soc1.conf (boot +7 s) | etsoc1-demo.service runs modprobe (boot +15 s) inferred |
| Commands sent to the card this boot (mgmt SQ0) | 0 on both cards | 3.25 million | 1.63 million |
| Driver error counters | et0: one PMIC correctable event; et1: none; no uncorrectable errors | MinionCe 23, SpCe 5 | SpCe 3 |
| Corrected PCIe errors at the card’s root port, this boot | et0 (00:01.0): about 695,000; et1 (00:01.1): 0 | 3 | 0 |
/var/log/kern.log | about 280 MB | 0.19 MB | 0.08 MB |
| Other ET driver on disk | older esperanto 0.20.0 DKMS package, built for 23 kernels, not loaded | none | none |
| Others on the machine | another user, logged in since 18 Sep; root GitHub Actions runner | our experiment queue; another user; a root GitHub Actions runner | our sampler; the etsoc1-demo and etsoc1-chatbot services |
Reproduced from source
Reproduced from source
We built the driver three ways on aifoundry2, in a scratch directory against the 7.0.0-31 headers, and installed nothing:
| Build | version | srcversion | Matches |
|---|---|---|---|
A: 09531e5c1 as it is | empty | 1383B256… | aifoundry1’s installed module |
B: 09531e5c1 with the Makefile from 78ed9b0d6 | 0.20.0 | 1383B256… | the minimal alternative in section 1 |
C: 353f20e | 0.20.0 | 47D26A30… | aifoundry2’s and aifoundry3’s installed module |
The code and data sections are identical in all three builds and in both installed modules. srcversion is a checksum of the driver’s .c and .h text, not of its Makefile. It differs only because commit 307e59c27 added a CentOS 9 condition, && (!defined(RHEL_MAJOR) || RHEL_MAJOR < 9), to five #if lines in et-soc1-pcie.c and et_pci_dev.h (and one in the loopback variant, which is not built). RHEL_MAJOR is not defined on Ubuntu, so the compiled code does not change. Nothing in the ET stack reads srcversion; DKMS uses it to compare builds.
How it got this way
- 21 Oct 2025.
09531e5c1introduces the typo. About 1.5 hours after it was committed,/usr/src/et-soc1-0.20.0was created on aifoundry1 (directory dated 19:40); its files match that commit. - 24 Oct 2025, 15:34. The source was registered again from the same tree inferred (file times fit
make dkms-remove; make dkms). A leftover build there, for kernel 6.14.0-33, already has the empty version, as every build from this directory must. - 26 Oct 2025. Upstream fix
78ed9b0d6. - Nov 2025 to 24 Sep 2026. Unattended kernel-header upgrades made DKMS rebuild the broken module from the stale copy 20 times, most recently for 7.0.0-34 on 24 September.
- 3 Jan 2026 and 21 Jul 2026. The fixed driver was built by hand on aifoundry1, in
/usr/src/et-platform/et-driver(for 6.14.0-37) and in/home/rehan/et-platform/et-driver(for 7.0.0-28). Both have version0.20.0, but neither was registered with DKMS, so each kernel upgrade left them behind. Whether they were ever loaded is not visible to us. - Mar 2026. aifoundry2 and aifoundry3 got DKMS from the fixed tree.
- 18 Sep 2026, 15:43. The current boot;
et_soc1was loaded at 15:47:30, and both cards’ nodes appeared within 2 seconds.
4. A separate problem: card et0’s PCIe link
This does not stop the cards from opening, and the module fix does not touch it. The root port above card et0 (00:01.0, the CPU’s x16 root port, leading to 01:00.0) had counted 695,018 corrected receiver errors (RxErr) and 5 bad DLLPs this boot at 13:34. They keep arriving at 1.0–1.4 per second while the card is idle, at random intervals. There are no non-fatal or fatal errors, and the link still runs at 16 GT/s x8. The errors are on the card-to-CPU direction: et0’s own counters are 0. For comparison, et1’s port (00:01.1) has 0, aifoundry2’s card has 3 for its whole boot under heavy load, and aifoundry3’s has 0. That is a bit error rate of at least about 10−11, roughly ten times the PCIe target: a marginal link.
Full detail: log growth, whether it varies between boots, the risk, and what to check
- It is what fills the logs. We sampled once a second for 97 seconds.
kern.lognever grew in a second without a new error. In 60 of the 70 seconds with errors it grew by exactly 527 bytes per error, the size of one four-line AER report; in the rest, the kernel’s rate limit (10 reports per 5 s) dropped some reports.sysloggrows the same way. That is about 50 MB a day in each file, which is whykern.logis about 280 MB and the journal now reaches back only to 22 September. The root file system has about 164 MB left (section 5). - It varies between boots inferred. From the compressed sizes of older logs, the stream was absent in the 30 August to 18 September boot and may have been present in a late-August boot. So one clean boot after a reseat is not proof of a fix.
- The risk inferred from the code. If the link ever produces a fatal error, the ET driver gives et0 up: it answers the kernel’s error recovery with DISCONNECT and tears down et0’s queues, so et0 stays unusable until it is probed again, in practice at a reboot.
dev_mngt_serviceandet-powertopopen every card’s management node at start-up, starting with et0, so they would then fail for et1 too. Programs built on aifoundry1 against its locally rebuiltlibdeviceLayer.acan be limited to one card withET_DEVICES=1, a local addition that is not in upstream et-platform.
What we suggest (root):
# limit the log to 10 reports an hour; the error counters keep counting; reverts at reboot
echo 3600000 > /sys/bus/pci/devices/0000:00:01.0/aer/correctable_ratelimit_interval_ms
# record the link detail only root can read (ASPM, link status, error masks, lane errors)
lspci -vvv -s 00:01.0; lspci -vvv -s 01:00.0
# watch the rate; the TOTAL_ERR_COR line should keep rising while kern.log stays quiet
cat /sys/bus/pci/devices/0000:00:01.0/aer_dev_correctableAt a maintenance window: after the module fix, note each card’s serial number (/dev/etN follows probe order). Then power off, reseat et0 and check its power connector. If the errors persist, swap et0 and et1: errors that follow the card mean the card, errors that stay on 00:01.0 mean the slot or board. Forcing that slot to Gen3 in the BIOS is a fallback. Judge the result by the hourly rate in aer_dev_correctable over at least two boots, not by the size of kern.log.
5. Also: a nearly full disk and a half-finished kernel update
Neither of these blocks the cards today, but both matter for the fix and for the next reboot.
The disk, the interrupted kernel update, and what to check
- The root file system is nearly full.
/has about 164 MB left (99% used, 13:57). The pool is 96% allocated, mostly by/home(421 GB) and/root(27 GB)./var/logtakes 268 MB on disk (170 MB of it the journal), and the AER reports from section 4 keep adding to it. aifoundry2 and aifoundry3 have hundreds of GB free. - The 24 September kernel update stopped partway. dpkg lists
linux-image-7.0.0-34-genericas half-configured (iF). Its lastdpkg.logentry, at 06:08:51, is the trigger that builds the initramfs and updates GRUB. After that:/boothasvmlinuz-7.0.0-34-genericbut noinitrd.img-7.0.0-34-generic, and the/boot/initrd.imglink points at the missing file;/boot/grub/grub.cfgwas last written on 5 September;/var/lib/dpkg/updatesstill holds entries from the interrupted run, and apt’shistory.loghas no end line for it.
grub.cfgwritten at 06:08 and 06:30. - What it means inferred. The next boot most likely still starts 7.0.0-31 from the September
grub.cfg. Booting 7.0.0-34 without an initramfs would not mount the ZFS root. A lack of space is a likely cause of the stop, but the installer’s output (/var/log/apt/term.log) is readable only by root and adm, so we could not confirm it. The module fix in section 1 already buildset_soc1for 7.0.0-34, so it does not depend on this.
What we suggest (root), after freeing some space:
df -h / # make room first
dpkg --configure -a # finishes the 7.0.0-34 install: initramfs and GRUB
ls -l /boot/initrd.img-7.0.0-34-generic # should now exist
dpkg -l linux-image-7.0.0-34-generic | tail -1 # should start with ii6. What we did not check or could not see without root
Caveats in full
- The cards’ firmware and state. No command has reached either card this boot, so their firmware versions, their device state and the runtime’s device-API check are unknown until the fix. The device nodes exist, and the driver creates them only after the firmware reports ready, so both cards were up when the module loaded inferred from the driver code. The firmware images in aifoundry1’s
/opt/etare a local build (acc7ed25d-dirty, 10 May 2026), and/opt/etswholds another set (d08fcdcb); aifoundry2 and aifoundry3 have353f20eimages. Images on disk need not match what is flashed: aifoundry2’s disk has BL 0.22.0 and minion 0.25.0 images, but its card reports BL 0.20.0 and minion 0.23.0. Which firmware aifoundry1’s cards run is unknown. - The kernel log.
dmesgand the journal are closed to us, so we did not see the probe-time messages, the one PMIC correctable event on et0, or when the RxErr stream started./var/log/dmesg(148 MB, readable by root and adm), written on 22 September byjournalctl --boot 0 --dmesg, probably holds this boot’s kernel log up to that day, including the 18 September probe inferred from its size. - What loads
et_soc1at boot +3 min 42 s. It is notmodules-load.dor udev. The candidates are the root GitHub Actions runner, a root cron job or a person inferred;/rootis not readable, so we could not read the runner’s workflow either. - Whether the hand-built fixed modules from January and July 2026 were ever loaded.
- The et0 link detail (ASPM, per-lane errors, error masks): PCIe config space beyond 64 bytes is root-only. Whether the card or the slot is at fault needs a reseat or a swap.
- Runtime parity after the fix. aifoundry1’s
libdeviceLayer.a(1 May 2026) andlibetrt.so(10 May 2026) were rebuilt from/home/rehan/et-platformand differ from aifoundry2’s. They are not part of this fault, but runtime programs on aifoundry1 will not run the same library binaries as on aifoundry2.
7. How the facts were gathered
Method, in full
- When. Friday 25 September 2026, 13:01–14:05 PDT. All three machines have been up since 18 September, 15:43.
- Read-only. As user
yaroslavvbover SSH, with no sudo: sysfs,/proc,modinfo,dkms status,lspcias a user, file listings and checksums, git in our clone of et-platform, and read-only git queries in the two et-platform checkouts on aifoundry1. Nothing was installed, and no module was loaded or unloaded. On aifoundry1 the only writes were two temporary decompressed copies of the module in our home directory, deleted at once. - No card was opened.
whoon aifoundry1 showed another user logged in at every check, from 13:01 to 14:05, so we opened no/dev/et*node there. On aifoundry2 and aifoundry3 our own experiment queues hold the cards, so we read only sysfs and files. - The error text is from our attempts on 22 September 2026. They failed before any command reached a card; the management queue counters are still 0.
- The rebuilds ran on aifoundry2 in a scratch directory (
make -C /lib/modules/7.0.0-31-generic/build M=<dir> modules). The installed modules of aifoundry1 and aifoundry2 were compared section by section. - Checked twice. Each conclusion was checked by a second, independent pass along three lines: the driver and its interface, the hardware, and how the driver was deployed. All three confirmed the cause. The second passes corrected details, not the conclusion: the verification commands, the log-flood link and the start date of the error stream. A last pass re-ran the cheap checks against this page (13:54–14:05) and corrected the kernel-headers claim for old kernels, the
#ifcount, the state of the 7.0.0-34 install and theET_DEVICESfilter in aifoundry1’s rebuilt deviceLayer.