Kathiravan Arumugam

Kathiravan Arumugam

AI/ML Engineer Trainee

Bangalore, IndiaArtificial Intelligence / Machine Learning
1roles
100skills
3education
3credentials

About

AI ML ENGINEER with hands on experience building production-oriented AI applicatiions using python, Langchain,Langgraph,FastAPI and pytorch . contributed to the development of generative AI ,retrival-Augmented generation(RAG),computervision,OCR and object detection/segmentation solutions during a one year engagement at Sony India Software Centre.Experienced in designing LLM powered applications,multiagent AI systems,and scalable AI services,Passionate about building production ready AI solutions that solve real world bussiness problem .

Experience

AI/ML ENGINEER TRAINEE - GEN AI & COMPUTER VISION

Sony India Software Centre

1 yr 1 mo · Bangalore Urban

AI/ML ENGINEER TRAINEE – Sony India Software Centre (SARD) • Contributed to Sony's SARD (Advanced R&D) team, developing AI solutions across Computer Vision and Generative AI for real-world applications. • Designed and optimized Retrieval-Augmented Generation (RAG) pipelines using LangChain and LangGraph, enabling AI agents to retrieve, reason over, and generate context-aware responses. • Developed and experimented with LLM-based workflows, prompt engineering, document processing, embeddings, vector retrieval, and multi-step AI agent orchestration. • Built and evaluated computer vision models for object detection and instance segmentation using PyTorch and OpenMMLab (MMDetection). • Performed dataset preparation, preprocessing, model training, hyperparameter tuning, and performance evaluation for deep learning models. • Explored model optimization and deployment techniques using ONNX and TensorRT to improve inference efficiency. • Collaborated with cross-functional engineers and researchers to analyze experimental results, troubleshoot model performance, and deliver production-oriented AI solutions.

Education

Nadar Saraswathi College of Engineering and Technology

BE, Electronics & Communication Engineering

Oct 2020 - May 2024

CGPA: 8.6

Professional course, Python Full stack development

Oct 2024 - May 2025

QSpiders - Software Testing Training Institute

Full stack development

Oct 2024 - May 2025

Skills

PythonLangChainLangGraphFastAPIPyTorchGenerative AIRetrieval-Augmented Generation (RAG)Computer VisionOCRObject DetectionImage SegmentationMMDetectionYOLOCNNsOpenCVHugging Face TransformersMultimodal LLMsGemma-3Prompt EngineeringModel Context Protocol (MCP)GPU TrainingmAPPrecision-Recall AnalysisConfidence Threshold TuningHyperparameter OptimizationNumPyPandasHTMLCSSSQLDockerCI/CD PipelinesGitHub ActionsModel DeploymentTensorRTONNXLinuxCondaGitGitHubAWS EC2AWS S3AWS Cloud DeploymentCUDAChromaDBGradioOpenRouterTavily Search APIPydanticPaddle OCRTrOCRAnalytical SkillsReasoning SkillsOrganization SkillsProblem SolvingCommunicationAgentic AI DevelopmentModel TrainingTransformer ModelsKnowledge Graph-Based Question AnsweringEnglishSQL Database AdministrationPython (Programming Language)Pandas (Software)TensorFloworacle SQLDjangoBootstrap (Framework)JavaScriptCascading Style Sheets (CSS)HTML5Embedded CArduinoC (Programming Language)Artificial Intelligence (AI)Software Architectural DesignTraining SystemsPost ProcessingConstrained OptimizationLearning EnvironmentCondosStabilizationFailure AnalysisPipelinesSystem MonitoringBilling ProcessComputational LinguisticsOptimization ModelsModel ValidationData PipelinesRobust OptimizationChatbot TestingWeb DevelopmentData AnalysisProgrammingLarge Language Models (LLM)Software DevelopmentAI&MLDeep LearningHugging Face

Projects

Interactive Multimodal RAG Chatbot with Automated PDF Reporting for Model Training Analysis

Designed and developed an interactive multimodal RAG chatbot using Gemma-3 (12B) to analyze model training reports and logs. Implemented a RAG pipeline over PDFs using MiniLM text embeddings and SigLIP image embeddings for visual-aware retrieval; extracted PDF images and generated contextual captions with a multimodal LLM. Built automated PDF reporting that consolidates YAML configurations, training logs, images, prompts, and LLM analysis, along with a Gradio UI for contextual Q&A, follow-up suggestions, and downloadable reports. Used Python, PyTorch, Hugging Face Transformers, CUDA, LangChain, ChromaDB, SigLIP, and Gradio.

AI-Powered Multi-Agent Deep Research System

Designed and developed a LangGraph-based multi-agent research system with agents for query clarification, research planning, web search, evidence synthesis, and report generation. Implemented hierarchical Supervisor, Research, Scoping, and Compression agents; integrated OpenRouter and Tavily Search API; built a FastAPI backend for REST APIs; and extended the platform with MCP support for local files, PDFs, and GitHub repositories. Used Python, LangGraph, LangChain, FastAPI, OpenRouter, Tavily Search, MCP, and Pydantic.

Real-Time Pothole Detection & Segmentation Pipeline

Developed a Mask R-CNN (ResNet-50 + FPN) instance-segmentation system for pothole detection. Built preprocessing, model-training, and evaluation pipelines using MMDetection; achieved 0.72 mAP; addressed varying lighting, road textures, and occlusions; and accelerated inference with TensorRT FP16 to approximately 9.58 ms latency. Used Python, PyTorch, MMDetection, OpenCV, Docker, and TensorRT.

Video-Based OCR Pipeline with Confidence Stabilization

Developed an end-to-end OCR system for extracting structured text from videos using frame sampling, ROI extraction, and preprocessing. Implemented confidence normalization, temporal stabilization, numeric validation, noise removal, and error-correction rules, and evaluated OCR performance across diverse video conditions. Used Python, OpenCV, Paddle OCR, TrOCR, NumPy, and Pandas.