PK Prabhat Kumar
Resume
Bengaluru, India 7+ YOE in AI / ML Open to Relocation (EU & Japan)

Prabhat Kumar

Machine Learning Engineer @ Roku • Former Researcher @ Bosch

Machine Learning Engineer & Researcher with 7+ years of experience across Computer Vision, Multi-modal Foundation Models, and Temporal Sequence Modeling. Passionate about bridging theoretical algorithmic formulation with robust, production-scale AI systems.

3 Published Patents (1 US, 2 IN) 🇪🇺 EU Blue Card Ready 🇯🇵 Japan HSP Eligible
Download CV
Prabhat Kumar
Shipped Systems & Research

Applied AI & Production Architectures

Selected machine learning architectures spanning multi-modal temporal reasoning, patented cross-camera perception, and edge-optimized robotics.

Roku March 2025 – Present

Multi-Modal Video Architecture & Temporal Modeling for Scene Break Detection

Scalable Long-Context Video Modeling & Boundary Localization
The Challenge

Long-form multi-modal scene break detection requiring high-throughput boundary localization across asynchronous video, audio, and subtitle streams where standard Transformers encounter quadratic memory scaling ($O(N^2)$).

Architecture & Method

Benchmarked linear-complexity State-Space Models (Mamba) and Graph Convolutional Networks (GCNs) against standard Transformers. Engineered self-supervised proxy tasks aligning visual (DINO), textual, and acoustic embeddings for precise boundary localization.

Scale & Impact

Scaled distributed feature extraction and validation across ultra-long video datasets using Ray, cutting processing bottlenecks and enabling localized structural reasoning.

PyTorch Ray State-Space Models (Mamba) DINO Transformers GCNs Distributed Scale
Bosch Research 2022 – 2025

Zero-Shot Perception Transfer & Cross-Camera Flow

Unsupervised Sensor Transfer & Dual-Level Domain Adaptation
The Challenge

Sensor recalibration and domain drift across heterogeneous multi-sensor vehicle and industrial cameras without expensive manual ground-truth annotation.

Architecture & Method

Coupled a UNIT-GAN translation network with a FlowFormer optical flow backbone to warp visual representations across pixel and latent feature levels, maximizing cross-domain geometric consistency.

Scale & Impact

Enabled unsupervised model and label transfer across heterogeneous sensors; directly produced 3 Published Patents (1 US Patent, 2 Indian Patents) in cross-camera vision.

Computer Vision FlowFormer UNIT-GAN Domain Adaptation PyTorch 3 Published Patents
Ola Electric 2021 – 2022

Low-Latency Autonomous Systems & Multi-Task Edge AI

Real-Time Embedded Multi-Task Perception for Autonomous Driving for Indian Roads
The Challenge

Joint perception for unstructured road environments under strict compute, memory, and millisecond latency constraints on embedded automotive hardware.

Architecture & Method

Engineered a unified multi-task deep network for simultaneous semantic segmentation and depth estimation trained with hybrid supervised and self-supervised geometric constraints.

Scale & Impact

Accelerated execution via TensorRT running within an onboard ROS/ROS2 robotics environment under sub-millisecond real-time constraints.

C++ ROS/ROS2 TensorRT Edge AI Multi-Task Learning Semantic Segmentation Depth Estimation
Career Trajectory

Work Experience

Over 7 years translating fundamental ML research into production algorithms across industry-leading teams.

Software Engineer, Machine Learning @ Roku

Dec 2025 – Present Bengaluru, India

Leading research and engineering in multi-modal video understanding architectures.

  • Researched multi-modal Scene Break Detection, scaling temporal sequence modeling by benchmarking linear-complexity State-Space Models (Mamba) and GCNs against standard Transformers.
  • Formulated self-supervised proxy tasks to align asynchronous cross-modal representations from visual (DINO), textual (Sentence Transformers), and acoustic foundation models for precise boundary localization.
  • Scaled distributed feature extraction and experimental validation across ultra-long-form video datasets using Ray, enabling high-throughput parallel processing and localized structural reasoning.

Researcher / Tech Lead & R&D Specialist @ Bosch Research / BGSW

Oct 2022 – Nov 2025 Bengaluru, India

Promoted to Tech Lead (Jul 2025); served as R&D Specialist (Oct 2022 – Jun 2025) driving autonomous perception.

  • Developed unsupervised cross-camera optical flow estimation by coupling a UNIT-GAN translation network with a FlowFormer backbone, enabling zero-shot model and label transfer across heterogeneous sensors.
  • Formulated dual-level domain adaptation objectives using predicted dense flow fields to warp visual representations at both pixel and latent feature levels, maximizing cross-domain geometric consistency.
  • Leveraged flow-based feature alignment to jointly regularize and improve representation capacity across source and target camera perception networks, resulting in 3 published patents.

Research Engineer I @ Ola Electric

Jul 2021 – Sep 2022 Bengaluru, India

Developed onboard computer vision and multi-task learning for Autonomous Driving for Indian Roads.

  • Developed a unified multi-task network for joint Semantic Segmentation and Depth Estimation tailored for unstructured road environments using hybrid supervised and self-supervised training.
  • Implemented domain adaptation pipelines and novel geometric and photometric data augmentations to minimize feature degradation against severe out-of-distribution driving conditions.
  • Evaluated model efficiency and real-world inference constraints by adapting neural architectures for low-latency execution via TensorRT within an onboard ROS robotics environment.

Student Researcher @ IISc & IIT Jodhpur

Aug 2018 – Jun 2021 Bengaluru & Delhi, India

Foundational academic research in self-supervised representation learning and deepfake forensics.

  • Indian Institute of Science (IISc): Formulated a self-supervised representation learning framework (SISL) to decouple content-independent camera signatures from image patches, training an anomaly localization network to identify splicing boundaries without pixel-level ground truth.
  • IIT Jodhpur: Developed a multi-stream deep neural network to isolate regional artifacts and spatial-temporal texture distortions from facial reenactment pipelines, mapping model robustness across progressive video compression codecs (published in WACV 2020).
Academic Foundations

Education

Degrees and graduate research in computer vision, deep learning, and data engineering.

2017 – 2019

M.Tech in Computer Science & Engineering (Data Engineering)

Indraprastha Institute of Information Technology Delhi (IIIT-Delhi)

Recognized institution (Anabin H+ accredited). Meets international equivalence standards for German EU Blue Card and Japan HSP.

2019 – 2021

Graduate Studies, Computer Vision & Deep Learning

Indian Institute of Science (IISc)

Research at Video Analytics Lab

2013 – 2017

B.Tech in Computer Science & Engineering

University of Delhi

Intellectual Property & Scholarly Work

Patents & Publications

Published patents in sensor perception and peer-reviewed research papers in top computer vision venues.

Published Patents (3)

Bosch Research
US IND DE Published 2025

System for Translation of Images Between Two Distinct Heterogeneous Cameras...

IND Published 2026

System for Adapting an AI Model Trained on a Source Camera and a Method Thereof

IND Published 2026

A Control Unit for Propogating Labels Across Different of Cameras using A Framework

Conference Papers

Google Scholar
CVPR Workshops (WMF) • 2022

SISL: Self-Supervised Image Signature Learning for Splicing Detection & Localization

Susmit Agrawal, Prabhat Kumar, Siddarth Seth, Toufiq Parag, Maneesh Singh, Venkatesh Babu

IEEE Winter Conference on Applications of Computer Vision (WACV) • 2020

Detecting Face2Face Facial Reenactment in Videos

Prabhat Kumar, Mayank Vatsa, Richa Singh

Technical Arsenal

Skills & Technologies

Core Research & Modeling

Model Architecture Design Generative AI Foundation Models Multi-modal Alignment Temporal Sequence Modeling Optical Flow Self-Supervised Learning Unsupervised Domain Adaptation Multi-Task Optimization Graph Neural Networks (GCNs)

Languages, Frameworks & Edge

Python PyTorch Ray TensorRT ROS / ROS 2 Distributed Training OpenCV CUDA
Perspectives & Culture

Beyond the Code

Global Collaboration & Engineering Mindset

Outside of algorithmic research and systems modeling, I thrive in collaborative, multi-cultural engineering teams. Having delivered across distributed teams in the US, Europe, and India, and having explored tech ecosystems firsthand in Stuttgart, Berlin & Yokohama, I appreciate environments that combine technical rigor with clear, empathetic communication.

Firsthand familiarity with Stuttgart, Berlin & Tokyo tech and cultural ecosystems.

Languages

English Full Professional / Fluent
Open to full-time international relocation