Video Platform
About 9 min read
← Back to work

ENGINEERING CASE STUDY

Real-Time Multi-Platform Video Integration

Unifying live and on-demand video across web, mobile, connected TV, and partner embeds — behind a single shared session contract.

ROLETechnical Program Manager, Video Platform
DURATION~6–9 months
TEAM~20 engineers across 4 pods
01CONTEXT & BUSINESS VALUE

Fragmented platforms, one fixed external deadline

The Problem

The organization operated multiple independently-built video experiences that had grown organically over several years: a browser-based web player, a native mobile SDK (maintained separately per platform), and a connected-TV (CTV) app built with an external vendor's involvement. Each surface had its own encoding pipeline, analytics schema, DRM handshake, and bug backlog. When a new business partnership required live video to be embeddable inside multiple external partner sites, engineering discovered there was no single, reusable “video integration layer” — every platform would need bespoke work, in parallel, against a fixed external launch date.

Core Problem

Fragmented, platform-specific codebases. No shared session or analytics contract. A firm external launch deadline.

Diagram showing legacy playback fragmentation: a new partner launch requirement branches into three separately implemented systems — legacy web player, legacy native SDK, and legacy CTV app — each with its own bespoke encoding, DRM logic, and analytics.
Legacy playback fragmentation — every partner launch was implemented three times over, once per platform.

Quantifying the Impact

The team framed success around three clear goals: reduce new-surface launch time, materially lower video-related Sev-1 incidents, and support large-scale concurrency without re-architecting later.

Onboarding Speed
New surfaces, fast
Reduce time to launch any new surface through a shared integration layer instead of one-off builds.
Reliability
Fewer Sev-1s
Shrink stream-start failures and licensing/token mismatches that were creating operational instability.
Scale
Large concurrency
Support a high-scale live audience at launch, with a path to grow further without another platform rewrite.

These directional goals — not just “make video better” — anchored every architecture and scope decision that followed.

My Role — Ownership Boundaries

I led the cross-functional delivery of the platform integration layer — the shared player core, the unified session/manifest service, the analytics event contract, and the CTV/partner-embed SDK surface — driving alignment across the engineering pods that built it, rather than authoring the implementation myself.

Drove

  • Architectural consensus on the shared session contract from discovery through sign-off.
  • Alignment across four engineering pods (web, native, CTV, partner/embed) on one contract and timeline.
  • Vendor escalation for licensing-latency risk, framing the business impact to unblock launch.
  • Program governance: cross-pod tracking, stakeholder dashboard, decision log.

Did Not Own

  • Hands-on implementation — I worked with each pod’s engineers rather than authoring the code.
  • Content encoding/transcoding pipeline — owned by media engineering.
  • CDN routing and edge caching — owned by infrastructure and reliability.
  • Ad insertion logic — owned by monetization, although I drove the integration requirements.

Making this delivery/ownership boundary explicit early — via a short scope document circulated to all contributing teams — prevented two common failure modes on cross-team platform work: ambiguity about who was accountable for implementation versus alignment, and decisions nobody actually owned.

02TECHNICAL ARCHITECTURE & CONSTRAINTS

One shared session contract, replacing four divergent paths

High-Level System Design

The core architectural decision was to introduce a unified video session service as the single source of truth between content origin and every player surface, replacing several parallel, platform-specific request paths.

Architecture diagram of the Unified Video Session Service: the content and encoding pipeline feeds manifests into the session service, which handles manifest resolution, license brokering, entitlement checks, and session token issuance, exposing a single contract to the web player, native SDK, CTV app, partner embed SDK, and internal QA harness.
The unified video session service sitting between the content pipeline and every platform surface.

Every surface spoke to one session contract: request a playback session → receive a signed manifest URL, a license endpoint, and a session ID used for all downstream analytics events. This collapsed several divergent request/response shapes into one versioned API, which is what made a much faster “new surface onboarding” goal achievable — a new platform only had to implement the client of a stable contract, not renegotiate playback logic from scratch.

Legacy Technical Debt We Faced

  • The web player had licensing logic hard-coded per content partner because it predated any concept of a shared entitlement service.
  • The native SDK used a bespoke analytics beacon format that other platforms could not parse, which made cross-platform funnel analysis impossible.
  • The CTV app had a slow, manual release process, meaning a broken contract change there wouldn’t surface for days rather than minutes.
  • None of the three had a shared definition of “playback start” — each platform measured it differently, which made reliability metrics internally inconsistent before any new work had even shipped.

Trade-offs Evaluated

OptionProsConsDecision
Rewrite all clients on a single cross-platform framework True code parity Infeasible within the launch timeline; sacrifices platform-specific performance tuning Rejected
Bespoke integration per partner, one at a time Fastest path to a single first partner Repeats the underlying fragmentation instead of fixing it Rejected
Shared backend session contract + thin, platform-native clients Each platform keeps native performance work; only the contract is unified; incremental adoption is possible Requires migrating existing clients onto the new contract, not just new ones Selected

The deciding factor was that a full client rewrite would have jeopardized the external launch date, while a backend-first contract let the fastest-to-migrate platform hit the deadline while the others migrated in parallel on a slightly longer runway — de-risking the hard date without abandoning the longer-term fix.

03PROGRAM STRATEGY & EXECUTION PLAN

Five phases, run in parallel across four pods

Lifecycle Phases

1

Discovery & Contract Design

Reverse-engineered the existing platform-specific request flows and defined the shared session API contract, with sign-off from all dependent teams on the interfaces they'd integrate against.

2

Foundation Build

Built the core session service (manifest resolution, license brokering, entitlement checks), stood up a unified analytics event schema, and migrated the web player first as the reference implementation.

3

Parallel Platform Migration

Native and CTV teams built against the now-stable contract concurrently, each on their own pod cadence.

4

Partner Embed SDK & Hardening

Built an iframe/postMessage-based embed SDK for external partner integrations, and load-tested against the concurrency target.

5

Launch & Stabilization

Phased rollout with incrementally increasing traffic, with rollback gates tied to error-rate thresholds at each stage.

Cross-Functional Alignment

With several engineering pods plus dependent media and monetization teams, I ran:

  • A weekly interface sync, kept short and scoped strictly to contract changes, so it didn't turn into a general status meeting.
  • A shared contract specification (versioned, with a changelog) as the actual source of truth, so “what does the session service return” was never a Slack-thread debate.
  • Platform-specific working sessions, run by each pod's own lead, for anything native to their stack — I stayed out of these unless a shared-contract change was needed, protecting the ownership boundary.

Governance & Tracking

  • A single cross-pod tracking epic with platform-specific sub-tracks, so leadership had one place to see aggregate progress.
  • A recurring metrics dashboard (build health, contract-conformance test pass rate per platform, load-test results) circulated to stakeholders on a regular cadence.
  • A decision log for every architecture trade-off, so later contributors could see why a choice was made, not just what was chosen.
04RISK, EDGE CASES & RESOLUTION

A third-party licensing bottleneck almost broke the launch

The Hidden Risk: License Latency Under Concurrent Load

During load testing at a meaningful fraction of our target concurrency, we discovered that the license-brokering step inside the session service — which synchronously called an upstream third-party licensing provider per session — began queueing under load. Latency for license issuance climbed sharply, well beyond what the playback flow could tolerate, threatening stream-start failures before we even reached the target concurrency level. This risk hadn't surfaced earlier because the legacy web player had never previously been tested anywhere near that scale.

This was the single biggest threat to the external launch date: it was discovered relatively late in the program, and it sat largely in a third-party provider’s infrastructure, outside our direct control.

Course Correction

  • Immediate: Introduced a local license-caching layer with short-lived reuse for identical entitlement/content combinations, substantially cutting redundant upstream calls in testing.
  • Structural: Moved the license-brokering call off the synchronous playback-start critical path where content policy allowed it, pre-fetching licenses ahead of anticipated high-demand live events.
  • Vendor escalation: Opened a formal capacity conversation with the licensing provider and secured additional headroom as a fallback safety net for launch week.
  • Re-ran load tests above the original target as a safety buffer and confirmed license latency stayed comfortably within acceptable bounds.
Sequence diagram of the license-caching flow: the client player requests a playback session from the session service, which checks the local license cache first. On a cache hit, the cached license is returned immediately. On a cache miss, the service makes a synchronous request to the upstream licensing provider, stores the new license with a short TTL for reuse, then returns the playback payload with URL, token, and license to the client.
The local license-caching layer added to the session service to absorb load on the upstream licensing provider.

Scope vs. Schedule Trade-offs

Descoped from launch

  • CTV deep-linking to specific live moments was descoped from launch to a later phase — a UX nicety rather than a reliability or contract requirement — freeing that pod to focus entirely on load and stability testing.
  • Partner embed theming/customization shipped with a single default look at launch, with a configurable theming layer added afterward, because building flexible theming under time pressure would have introduced new risk.

Never cut

  • Load testing and hardening work that validated the core reliability commitment.
  • The shared session contract itself, which was the foundation for the rest of the program.

The operating principle was simple: cut scope that adds new surface area under pressure, but never cut the testing and hardening work that validates the core reliability commitment.

05RESULTS & ENGINEERING RETROSPECTIVE

From fragmented builds to a reusable platform layer

Data-Backed Outcomes (directional, measured shortly after General Availability)

Onboarding Speed
Substantially faster
New platform integration time dropped from the historical baseline, beating the onboarding-speed target.
High-Severity Incidents
Cut by more than half
Fell versus the prior baseline period, exceeding the original reduction goal.
Peak Concurrency
Well above target
Sustained during a high-traffic live event, with low stream-start latency and no licensing failures.
Analytics Parity
Unified
A single “playback start” definition eliminated cross-dashboard reporting discrepancies.

Post-Mortem — What Went Wrong

  1. The licensing load risk should have been caught earlier in Discovery. The architecture review assumed the legacy system’s behavior would hold under scale rather than treating “untested at scale” as a risk in its own right. Retrospective action: future migrations of legacy systems now require explicitly load-testing the old system before designing the new one.
  2. The CTV vendor's slow release process was underestimated as a schedule risk. Sprint planning didn't fully account for multi-day release/review cycles, which compressed real iteration time. Vendor release-cycle constraints are now built into program timelines explicitly rather than treated as a footnote.
  3. Cross-pod governance started a bit later than ideal, leading to some early duplicated contract-design effort between two pods before the shared sync existed. Lesson: stand up lightweight cross-pod governance from day one, even before there's much to synchronize on.

Future State & Handoff

  • Next phase (already scoped): a theming API for partner embeds, CTV deep-linking, and extending the session service to support server-side ad insertion natively.
  • Ownership handoff: the session service moved from the project team to a permanent platform on-call rotation, with the contract specification and decision log handed off as living documentation.
  • Scaling roadmap: a further concurrency scaling target is planned ahead of the next major live-event campaign, with the license-caching layer identified as the next component to pressure-test first.

Closing Reflection

The work became more than a launch fix. It created a reusable platform foundation that made future video integrations faster, safer, and easier to scale.