Resolve memory leak for server-initiated room closure - #1473
alan-george-lk wants to merge 3 commits into
Conversation
Changeset ✓This PR includes a changeset covering all affected packages:
|
There was a problem hiding this comment.
Devin Review found 1 potential issue.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
There was a problem hiding this comment.
🟡 Failed unpublish sends contradictory binding events
When sender removal fails, unpublish_track now emits LocalTrackUnpublished but returns an error. The FFI unpublish request forwards that event, then reports failure to its caller.
(Refers to this code)
Learn more
The room emits LocalTrackUnpublished from this callback. The FFI event handler forwards it and removes the publication lookup entry. Meanwhile, the FFI request treats unpublish_track's returned error as a failed operation and sends an error callback. A closed peer connection triggers this mismatch even though the local track has been detached. Clients can receive a publication-removed event alongside a failed unpublish request.
Example: A participant unpublishes camera TR_1 while a server deletion closes its connection. Sender removal fails, so the SDK detaches TR_1 and sends LocalTrackUnpublished; the FFI request then reports an error for the same unpublish.
Recommended fix: Decide how the API represents completed local cleanup after sender-removal failure. Coordinate the unpublish_track return value with FFI's callback and event forwarding so the same operation has one consistent outcome; preserve the underlying removal error through logging or a distinct diagnostic if it cannot be returned as the operation result.
Was this helpful? React with 👍 or 👎 to provide feedback.
This fixes a fairly hefty memory leak reported by a customer. Their use case is remotely killing a room via server API, which kills the client room locally, and joining another room in the same client process.
Testing summary below, where a "cycle" is a single process that:
livekit --dev room delete <room name>Previously,
unpublish_trackcalledremove_track(sender)?first. When a remote room deletion had already closed the connection, that call errored and returned early, leaving thepublication → track → transceivergraph attached, which could retain the room session.The patch:
remove_trackresult instead of returning immediately.The added reconnection test publishes a local video track, deletes the room remotely, then verifies both that the publication loses its track and that the room session can actually drop. That directly demonstrates the retained-reference lifecycle issue is resolved.
I cut a specific C++ SDK branch build against this fix and the customer confirmed this resolved the problem.