First draft, written from the build notes.
What it is
Jellyfin is the media server in my house. Its web client works, but it feels like a settings page with posters. Azushi replaces that client with an interface modelled on the Apple TV app: a featured title across the top, Continue Watching merged across films and shows, details that rise as a sheet, and a player of its own.
It is one plugin. Install it on the server and every browser that opens Jellyfin gets Azushi; the admin dashboard is left as stock Jellyfin.

Getting in
Jellyfin 12 has no way for a plugin to reach the web client. The only sanctioned hook, the branding options, carries CSS and nothing else, and CSS can't carry an application.
What it does have is IPluginServiceRegistrator, which hands a plugin the server's real service collection. So Azushi registers its own ASP.NET Core IStartupFilter: a small middleware that rewrites jellyfin-web's index.html on its way out and adds the Azushi bundle. No third-party plugin, no patching of the server at runtime.
One trap only a real browser shows: a startup filter runs after Jellyfin's response compression, so the middleware was reading gzip as text and the page rendered as garbage. curl never showed it, because curl doesn't ask for compression. The middleware now removes Accept-Encoding from that one request so nothing compresses it, and drops the old ETag once the body has changed.
An overlay, not a fork
Azushi doesn't replace jellyfin-web; it sits on top of it and takes over every route except the dashboard. Playback, sign-in, sessions and casting all stay Jellyfin's, so they keep working exactly as Jellyfin's do, and Azushi only has to be good at what you see.
The player is the clearest case. Its controls are Azushi's own, in the style of tvOS, but every play, pause and seek drives jellyfin-web's playback engine underneath rather than a second engine competing with it.

Hearing the intro
Skip Intro needs to know where every intro is. Azushi works it out from the audio, with no external tools: the server's own ffmpeg decodes each episode to 16 kHz mono, and each frame becomes a 32-bit fingerprint whose bits record how the energy in 33 bands changed. Two episodes share an intro where their fingerprints line up.
The first version found 11 of 16 test intros. All five misses were the same problem: two episodes put the intro at different offsets, so one's frames landed between the other's. Widening each frame from 128 to 512 milliseconds while keeping the 64-millisecond step fixed all five.
The rest is structure. A short match seeds a candidate and grows outward while the audio still agrees; each episode is compared with its next two neighbours and the largest agreeing group wins; the results go into Jellyfin's own segment store, so Skip Intro works in every Jellyfin client, not just Azushi.
Three days of decoding
On my server's library, about 28,000 episodes, detection ran ffmpeg around the clock for three days: 86,372 seconds of decoding on one day alone.
The cause was in how Jellyfin 12.1 schedules this work. Every twelve hours it asks every segment provider about every item, and an empty answer deletes what was stored. Azushi had relied on a cache that a restart wiped, so each pass started from nothing.
The fix is a ledger. For each season it keeps a fingerprint of the files and the segments found, including "none". A season that hasn't changed is a dictionary lookup; one that already has segments from elsewhere is adopted rather than decoded; only a changed season is analysed again.

Shipping it
Changes reach main only through squash-merged pull requests with Conventional Commit titles, and Release Please turns them into versions. Updates arrive through Jellyfin's own plugin page: the plugin serves a private repository on the server's loopback address, backed by GitHub Releases, so the source stays private and the server updates like it would from any public catalogue.
What I would change
For Manraj to write.