Lifecycle & Memory Safety¶
Rules for Thing references, shared resources, and connection lifecycle.
Part of compliance-rules.
RULE-002: No Direct Thing References Across Async¶
Severity: Critical
Description:
Framework services (schedulers, callbacks, long-lived closures) must not hold direct Thing references. Store Guid IDs and resolve from DeviceTreeRegistry at execution time. Direct references prevent garbage collection when Things are removed/reconnected.
Short-lived command closures that use the official Thing.ExecuteCommand(...) pattern are allowed. They are scoped to the command task and are released on command completion/cancellation with the owning Thing tree. Do not report those as RULE-002 unless the closure is retained by a long-lived root such as a scheduler, timer, singleton/shared resource, static field, persistent event subscription, or a pending task that is not cancelled by Thing/protocol teardown.
Detection pattern:
- Fields of type
Thingor subclass in services, schedulers, or singletons - Closures capturing Thing instances in long-lived
asynccallbacks, timers, schedulers, subscriptions, or queues that outlive Thing/protocol teardown Dictionary<*, Thing>orList<Thing>in long-lived objects- Do not flag
ExecuteCommand(...)command-handler callbacks solely because they capture the command target Thing or its parent for the duration of the command
Correct pattern:
// Store ID, resolve at execution time
private readonly Guid _thingId;
public async Task Execute(DeviceTreeRegistry registry, CancellationToken ct)
{
var thing = registry.Resolve(_thingId);
if (thing is null) return; // Thing was removed
// ... use thing
}
Violation example:
// WRONG: holds reference across async boundary
private readonly PanelDevice _panel;
public async Task DoWork()
{
await Task.Delay(5000);
await _panel.SomeMethod(); // _panel may be stale
}
Allowed example:
// OK: framework command supervision pattern; closure lifetime is the command task.
return door.ExecuteCommand(
DoorState.Unsecured,
async () => await protocol.SendCommand(module.Address, 2, ct),
protocol,
e => e.Address == module.Address && e.EventType == AcmeEventType.StrikeReleased,
onSuccess: async e => await door.UnsecuredRemote(e.Timestamp, command.OperatorIdentity),
ct: ct);
RULE-017: Shared Resource Lifecycle Must Be Complete¶
Severity: High
Description:
ISharedResource implementations (SDK processes, connection pools, gRPC channels shared across multiple Things) must implement complete lifecycle:
Start()- initialization + wait for ready signal with timeoutStop()- graceful shutdown with timeout, kill fallback- External processes must wait for "ready" signal before Start returns
Why this matters:
- Without ready check: Things use resource before initialization completes → crashes, races, lost data
- Without timeout: hung process blocks adapter startup indefinitely
- Without Stop: zombie processes, leaked sockets, file handles
- Framework starts Things only after
Start()returns - timing matters
Detection pattern:
ISharedResourceimplementation missingStart()orStop()Start()returns immediately without waiting for resource readiness- External
Process.Start()without subsequent ready check (port probe, log scan, IPC handshake) Stop()without timeout (or no kill fallback when graceful shutdown fails)- No cancellation token propagation in lifecycle methods
Correct pattern:
internal sealed class VendorSdkProcess : ISharedResource
{
private Process? _process;
public async Task Start(CancellationToken ct)
{
_process = Process.Start(new ProcessStartInfo
{
FileName = "vendor-sdk.exe",
RedirectStandardOutput = true,
});
// Wait for ready signal with timeout
using var timeout = CancellationTokenSource.CreateLinkedTokenSource(ct);
timeout.CancelAfter(TimeSpan.FromSeconds(30));
await WaitForReady(_process, timeout.Token);
}
public async Task Stop(CancellationToken ct)
{
if (_process is null) return;
// Graceful shutdown
_process.StandardInput.WriteLine("shutdown");
// Timeout + kill fallback
if (!_process.WaitForExit(5000))
_process.Kill(entireProcessTree: true);
await _process.WaitForExitAsync(ct);
_process.Dispose();
}
private static async Task WaitForReady(Process process, CancellationToken ct)
{
while (!ct.IsCancellationRequested)
{
var line = await process.StandardOutput.ReadLineAsync(ct);
if (line?.Contains("READY") == true) return;
}
throw new OperationCanceledException("SDK process did not signal ready");
}
}
Violation examples:
// WRONG #1: no ready check
public Task Start(CancellationToken ct)
{
_process = Process.Start("vendor-sdk.exe");
return Task.CompletedTask; // Things may use SDK before it's initialized
}
// WRONG #2: no timeout, hangs forever
public async Task Start(CancellationToken ct)
{
_process = Process.Start("vendor-sdk.exe");
while (!await IsReady()) // no timeout, no CT propagation
await Task.Delay(100);
}
// WRONG #3: no kill fallback
public async Task Stop(CancellationToken ct)
{
_process?.CloseMainWindow();
await _process.WaitForExitAsync(ct); // hangs if process ignores close signal
}
Framework responsibility:
SharedResourceManagercallsStart()before any Thing that depends on the resourceStop()called during adapter shutdown after all dependent Things stop- Failure in
Start()propagates as adapter startup failure - log clearly
Verification:
- Kill SDK process externally → adapter detects failure, retries via reconnect logic
- Cancel adapter during startup → Start returns within timeout, no leaked process
- Restart adapter multiple times → no zombie processes, no port conflicts
RULE-018: Connection Must Be Fully Ready Before Returning True¶
Severity: Critical
Description:
IDeviceConnection.Connect() must return true ONLY after the connection is fully usable. Premature true causes framework to start polling, commands, and event subscriptions against an unready connection - resulting in lost commands, dropped events, and false-positive status reports.
What "fully ready" means:
- TCP/transport connected
- Authentication completed
- Protocol initialized (handshake, version exchange, capability negotiation)
- Address binding/registration done (if protocol requires)
- Event stream/polling readiness ensured (subscriptions registered, queues drained)
Failure semantics:
- Any step fails → return
false→ framework retries with exponential backoff - Return
trueonly when ALL steps succeed - Do NOT call
SignalConnected()fromIDeviceConnection- framework owns that
Detection pattern:
Connect()returns true after TCP connect without authentication- Missing protocol initialization (handshake, login response check) before true return
- Authentication failure swallowed (catch + return true)
SignalConnected()called insideIDeviceConnection.Connect()implementation- Polling/event subscription not verified before returning true
Correct pattern:
public async Task<bool> Connect(CancellationToken ct)
{
try
{
// 1. Transport
await _client.ConnectAsync(_panel.IpAddress, _panel.Port, ct);
// 2. Authentication
if (!await Authenticate(_panel.Password, ct))
{
_logger.LogWarning("Authentication failed for panel {Panel}", _panel.IpAddress);
return false;
}
// 3. Protocol initialization
var version = await _protocol.GetVersion(ct);
if (!version.IsSupported)
{
_logger.LogError("Unsupported panel version: {Version}", version);
return false;
}
// 4. Address binding (if needed)
await BindRequiredAddresses(ct);
// 5. Event subscription
await _protocol.SubscribeEvents(ct);
return true; // fully ready
}
catch (OperationCanceledException)
{
throw; // let CT propagate
}
catch (Exception ex)
{
_logger.LogWarning(ex, "Connect failed, will retry");
return false; // framework retries with backoff
}
}
Violation examples:
// WRONG #1: returns true after TCP, before authentication
public async Task<bool> Connect(CancellationToken ct)
{
await _client.ConnectAsync(_panel.IpAddress, _panel.Port, ct);
return true; // framework starts polling, but we're not authenticated yet
}
// WRONG #2: swallows authentication failure
public async Task<bool> Connect(CancellationToken ct)
{
await _client.ConnectAsync(_panel.IpAddress, _panel.Port, ct);
try { await Authenticate(ct); } catch { /* ignored */ }
return true; // claims success even on auth failure
}
// WRONG #3: missing protocol initialization
public async Task<bool> Connect(CancellationToken ct)
{
await _client.ConnectAsync(ct);
await Authenticate(ct);
return true; // no handshake, version check, or event subscription
}
// WRONG #4: calling SignalConnected in IDeviceConnection
public async Task<bool> Connect(CancellationToken ct)
{
await DoConnect(ct);
_connectionFunction.SignalConnected(); // framework does this, not adapter
return true;
}
Framework guarantees:
- After
Connect()returns true: framework callsOnConnected()onIConnectionAwareThings in subtree - StatusPoll, HistoryPoll, TimeSync schedulers activate for this Thing
- Commands routed to this Thing become eligible for dispatch
- Event interceptor active for
WaitForEventpatterns
Verification:
- Disconnect device mid-connect → next
Connect()retries from step 1 - Wrong credentials → returns false, framework retries (eventually backs off)
- Network glitch during handshake → returns false, no partial state leaked
- Connect returns true → first command/poll immediately succeeds
RULE-052: Time Synchronization Must Be Enabled When the Protocol Can Set the Device Clock¶
Severity: High
Description: When the vendor protocol documents a set-clock/set-time command, the adapter must wire time synchronization completely. Device clocks drift; an unsynced device stamps history events with wrong timestamps, which corrupts the audit trail (RULE-008 depends on device timestamps being trustworthy).
Time sync has two mandatory halves, and each fails differently when missing:
capabilities.time_sync: trueinadapter-registration.yaml— without it the generator never adds theTimeSynchronizationfunction to the root device type. Silent failure: compiles clean, the device clock is simply never synced.Capability.ITimeSynchronizedimplemented on the root Thing's partial —TimeSynchronizationFunctioncasts its owner to the interface in its constructor, so a missing implementation with the capability enabled fails loudly at startup.
The dangerous half is #1: an adapter that implements ITimeSynchronized but omits the capability flag looks complete in source and never syncs.
Detection pattern:
- Protocol layer (
Protocol*.cs, vendor docs) documents a clock-set command, butadapter-registration.yamlhas nocapabilities.time_sync: true— violation. ITimeSynchronizedimplemented anywhere without the capability flag (or vice versa) — violation; the two must appear together. The build enforces this clause as PQC152, in both directions.- Clock-set protocol method exists but is called from nowhere except command handlers — the periodic sync path is missing.
Correct pattern:
public partial class AcmePanel : Capability.ITimeSynchronized
{
/// <inheritdoc/>
public Task SynchronizeTime(TimeProvider timeProvider) =>
Protocol.SetClock(timeProvider.GetLocalNow());
}
Acceptable exception (must be documented): No time sync is acceptable only when the protocol exposes no clock-set command, or the device manages its own clock (e.g. NTP) — documented in the adapter's design-notes document under Known Implementation Gaps (or an equivalent design note).
RULE-057: Set the Device Clock From the Wall Clock, Never a Re-Projected Local Time¶
Severity: High
Description:
SynchronizeTime(TimeProvider timeProvider) receives a timezone-aware provider whose LocalTimeZone is the device's configured zone (TimeSyncScheduler builds it from the TimeSynchronizationFunction.Timezone property). timeProvider.GetLocalNow() therefore returns a DateTimeOffset whose components already carry the correct device wall clock, e.g. 11:22 +02:00.
The trap is converting that DateTimeOffset to a DateTime with .LocalDateTime: that property re-projects the instant onto the host process zone — which in a container is UTC — yielding 09:22. The panel is a wall-clock device (no timezone concept), so it gets set off by the container↔device offset, and every event it stamps afterward is shifted by that offset. The failure is silent: it compiles, syncs "successfully", and only surfaces as audit timestamps that are hours wrong.
DateTimeOffset.DateTime would give the right value but is banned by analyzer policy because it drops Kind.
The fix is to keep the offset, not to convert it. Make the protocol-layer set-clock method accept a DateTimeOffset and read its components (.Hour, .Day, .Month, .Year, .Offset, …) — those are already in the device zone. Then SynchronizeTime is a one-liner that passes GetLocalNow() straight through.
Detection pattern:
GetLocalNow().LocalDateTimeanywhere in aSynchronizeTime/ set-clock path — violation.- A set-clock protocol method typed
DateTimefed from aTimeProvider— smell; retype the boundaryDateTimeOffset. GetLocalNow()passed throughTimeZoneInfo.ConvertTimeFromUtc(...),new DateTime(now.Year, …), or any other zone math on the way to the wire — flag; the only sanctioned form is passingGetLocalNow()straight through.- A time-sync or device-event-timestamp path reaching for
TimeProvider.System(or aTimeProviderfield defaulted to it) instead of the zone-aware provider the framework hands toSynchronizeTime— flag; it stamps the host zone, not the device zone. For an event that arrives with no device timestamp, useDeviceTimestamp.UtcNow(absolute receipt time), never a fabricated wall clock.
Correct pattern:
// Thing
public Task SynchronizeTime(TimeProvider timeProvider) =>
Protocol.SetClock(timeProvider.GetLocalNow()); // pass the offset straight through
// Protocol — consumes the offset, reads components in the device zone
public Task SetClock(DateTimeOffset time) =>
SendClockFrame(time.Year, time.Month, time.Day, time.Hour, time.Minute, time.Second);
timeProvider.GetLocalNow() is the only value that may reach the wire. Any transformation of it raises a flag — .LocalDateTime and .DateTime are the two traps above, but TimeZoneInfo.ConvertTimeFromUtc(...), new DateTime(now.Year, now.Month, …), or any other hand-rolled zone math is equally a signal that a DateTime-typed boundary is being fought instead of fixed. Retype the boundary to DateTimeOffset; do not launder the value. (Reading the device clock back for a drift check is different: wrap the device's own naive reading — new DateTimeOffset(deviceReading, now.Offset) — and compare it to the untransformed now.)
Violation example:
// WRONG: .LocalDateTime re-projects 11:22 +02:00 onto the container zone (UTC) → 09:22
public Task SynchronizeTime(TimeProvider timeProvider)
{
var localNow = DateTime.SpecifyKind(timeProvider.GetLocalNow().LocalDateTime, DateTimeKind.Unspecified);
return Protocol.SetTime(localNow); // panel clock ends up off by the container↔device offset
}
The build enforces the first detection clause as PQC257: a GetLocalNow() call whose result is read through .LocalDateTime is reported on every build, local and CI. The other clauses — zone math on the way to the wire, a DateTime-typed set-clock boundary, a path reaching for TimeProvider.System — stay with review.
Verification:
- Run the adapter in a UTC container against a device in a non-UTC zone; the device clock must match the device wall time, not container UTC.
- The zone used to set the clock (this rule) must equal the zone used to read event wall clocks back (
EventBuilderExtensions.Atresolves it from the sameTimezone); a mismatch here is exactly what shifts audit timestamps.
RULE-058: An Authentication Verdict Must Come From the Device, Never From a Transport Failure¶
Severity: High
Description:
When a Thing declares the AuthenticatedProtocol function, its connect path carries a second, operator-facing verdict alongside true/false: why the attempt failed. protocol.authentication.failed and protocol.handshake.failed tell the operator that retrying is pointless until someone changes a credential or a certificate — that is the entire value of the function. A transport failure reported as an authentication failure sends the operator to fix a password while a cable is unplugged, and a stale failure state left behind after a successful login makes a healthy panel look permanently rejected.
What the rule requires:
- Every exit of the connect path sets the function exactly once.
SetAuthenticationFailed()/SetHandshakeFailed()only when the device answered and refused — a vendor error code, a typed SDK exception, or a protocol response says so.- Every other failure — timeout, refused port, DNS, socket reset, unclassified exception — uses
SetDefault(), which clears the slot without asserting anything. - The success path calls
SetAuthenticated(), clearing any earlier failure. - A session rejected or expired mid-run sets
SetAuthenticationFailed()before the disconnect, so the retry loop stays explained.
Declaring the function at all is a claim that the adapter can make this distinction. If the protocol collapses every failure into one opaque timeout, do not declare AuthenticatedProtocol — an always-empty slot is honest, a guessed one is not.
Detection pattern:
- A
catch (Exception)/ barecatchthat callsSetAuthenticationFailed()orSetHandshakeFailed() - A failure exit of the connect path that sets no state at all, leaving the previous verdict standing
SetAuthenticated()called before the login/handshake result was checked (see RULE-018 — the same premature-success mistake)- No
SetAuthenticated()on the success path, so a recovered device keeps a failure status forever - A mid-run session rejection routed only to disconnect, with the protocol slot untouched
AuthenticatedProtocoldeclared on a Thing whose protocol exposes no distinguishable authentication error
Correct pattern:
private async Task<bool> Connect(Panel panel, CancellationToken ct)
{
try
{
await _client.Open(_endpoint, ct);
await _client.Login(_userName, _password, ct);
await panel.Functions.AuthenticatedProtocol.SetAuthenticated(ct);
return true;
}
catch (BadCredentialsException)
{
// device answered and refused - retrying will not help
await panel.Functions.AuthenticatedProtocol.SetAuthenticationFailed(ct);
return false;
}
catch (SecureChannelException ex)
{
// failure is known to be inside secure-session setup, but not which credential
_logger.LogWarning(ex, "Secure handshake failed");
await panel.Functions.AuthenticatedProtocol.SetHandshakeFailed(ct);
return false;
}
catch (Exception ex)
{
// transport level - no authentication verdict to report
_logger.LogWarning(ex, "Connect failed, will retry");
await panel.Functions.AuthenticatedProtocol.SetDefault(ct);
return false;
}
}
Violation examples:
// WRONG #1: every failure becomes an authentication failure
catch (Exception ex)
{
await panel.Functions.AuthenticatedProtocol.SetAuthenticationFailed(ct);
return false; // unplugged cable now reads as "wrong password"
}
// WRONG #2: transport failure leaves the previous verdict standing
catch (SocketException)
{
return false; // stale protocol.authentication.failed survives the network outage
}
// WRONG #3: success never clears the failure
await _client.Login(_userName, _password, ct);
return true; // no SetAuthenticated - panel stays red after credentials are fixed
// WRONG #4: verdict asserted before the device answered
await _client.Open(_endpoint, ct);
await panel.Functions.AuthenticatedProtocol.SetAuthenticated(ct);
await _client.Login(_userName, _password, ct); // may still be refused
Interaction with polling:
The state is produced by the connect and session path, not by a device register. When the Thing also declares StatusPoll, the state must be declared as excluded from poll coverage rather than left silently uncovered:
[PollExcludes(typeof(AuthenticatedProtocolState), "Protocol authentication is reported by the connection handshake path.")]
Verification:
- Wrong credentials →
protocol.authentication.failed, and it persists across retries - Correct credentials afterwards → status clears on the first successful login
- Pull the network cable → slot goes empty (
SetDefault), onlyconnection.*reports the outage - Point the adapter at a closed port → no protocol status appears at all
- Revoke the session on a running device → failure status appears without waiting for the next connect