Manmohan Vishwakarma

Manmohan Vishwakarma

Deep Learning and Data Enthusiast

Odisha, IndiaArtificial Intelligence / Computer Vision
2roles
60skills
4education
2credentials

About

AI/ML engineer with strong focus on computer vision and video understanding, experienced in building end-to-end, production-oriented ML systems. Hands-on experience in action recognition, spatio-temporal modeling, dataset engineering (AVA-style formats), and deep learning using PyTorch and TensorFlow. Actively working on open-source and research-oriented projects involving video analytics, surveillance intelligence, and scalable ML pipelines. Comfortable owning the full lifecycle—from data curation and model training to deployment-ready tooling—and continuously strengthening foundations in DSA and system-level thinking to operate effectively at scale.

Experience

AI/ML Engineer

Deepnet Analytics

8 mos · New Delhi, Delhi, India

AI / Computer Vision Intern

DeepNet Analytics

June 2025 - Jan 2026

Co-developed OpenAVA-Lab, an open-source AVA-style spatio-temporal dataset creation platform reviewed at AAAI 2025. Built an RF-DETR and KalmanSORT automated proposal pipeline at 25+ FPS, with 7.2x throughput gain and 87% reduction in active annotation time. Exposed orchestration through FastAPI REST APIs and deployed on AWS EC2, Lambda, and RDS using Docker and PostgreSQL. Implemented YOLOv9 PPE detection achieving 90% mAP on a custom 8-class industrial dataset.

Education

National Institute of Technology , Patna

Bachelor of Technology - BTech,Electronics and Communications Engineering

Jul 2023 - Jul 2027

St. Gregorios High School

High school

[object Object]

DR.A.N.Khosla Dav Public School

plus 2

[object Object]

National Institute of Technology, Patna

B. Tech, Electronics & Communication Engineering

2023-2027

CGPA: 8.21/10. Co-author of OpenAVA-Lab, AVA-style Spatio-Temporal Dataset Platform (reviewed at AAAI 2025; positive reviewer feedback).

Skills

DockerAI AgentsArtificial Intelligence (AI)Machine LearningPython (Programming Language)MongoDBFastAPIJavaC (Programming Language)C++HuggingfaceLarge Language Model Operations (LLMOps)Large Language Models (LLM)LangChainNatural Language Processing (NLP)DjangoTkinterKaggleOpenCVJinjaFlaskKerasPandasPyTorchTensorFlowCascading Style Sheets (CSS)NumPyPythonSQLYOLOv9RF-DETRLW-DETRBERTTransformersMMAction2MediaPipeMTCNNDeepfake DetectionLiveness DetectionRAGLLMsAWSEC2LambdaRDSPostgreSQLRedisCVATCI/CDGitHub ActionsREST APIsMicroservicesApache KafkaNginxGitLangGraphChromaDBSciPyStreamlitLinux

Projects

HAWK-1 - Industrial Safety AI System

Startup MVP. Built an AI-powered PPE compliance system achieving 95% mAP across 8 safety-equipment classes with end-to-end inference latency under 60 ms. Processed 10,000+ annotated frames and architected an event-driven Kafka alert system for concurrent camera streams. Deployed Dockerized microservices on AWS GPU instances.

Pebbles - Agentic Business Analytics Dashboard

Built a LangGraph multi-agent analytics platform with a RAG pipeline achieving 80% query accuracy and an asynchronous FastAPI REST layer for conversational business-metrics queries.

AuthKYC - Presentation Attack Detection for Bank KYC

Built an end-to-end 4-stage cascading presentation-attack-detection security waterfall achieving 95.2% accuracy and 0% false-positive rate on a mixed-domain evaluation set. Designed a Frequency-Temporal Cross-Attention architecture and implemented CHROM rPPG liveness, PRNU sensor-noise forensics, and Moiré FFT replay detection.