{"domain":"https://www.chandraprakashs.im","agentReady":true,"specification":"https://www.chandraprakashs.im/llms.txt","knowledgeBase":"https://www.chandraprakashs.im/llms-full.txt","sitemap":"https://www.chandraprakashs.im/sitemap.xml","profile":{"name":"Chandra Prakash S","role":"Software Engineer","specialty":"AI Systems & Full Stack","phone":"+91-7780711026","tagline":"Engineering Intelligence. Building for People.","bio":"Software Engineer building production-grade AI systems, Conversational AI, Generative AI tools, multi-provider LLM orchestration, and low-latency infrastructure born from real friction.","aboutNarrative":"I build software because the problem deserves a solution. Almost every project I've built—from ClarityUX and AI Workforce Planning to Sunflower AI—started from real friction rather than theoretical exercises. I focus on fewer moving parts, clean mental models, and invisible engineering that leaves users feeling effortless.","philosophy":"Technology is a means; the product is the outcome. Simplicity is sophistication. Engineering should disappear—users shouldn't notice architecture or AI complexity, only that it feels effortless.","education":{"institution":"Dr. M.G.R Educational and Research Institute","degree":"Master of Computer Applications (CGPA: 8.05)","secondaryDegree":"SRM Institute of Science and Technology - B.Sc. Computer Science (CGPA: 8.6)","description":"Specialized in advanced computer science fundamentals, machine learning architecture, distributed backend systems, and software design."},"location":"Bengaluru, India","email":"chandra.1997@hotmail.com","status":"Open to Opportunities","deployedProjectsCount":"5+"},"socials":[{"name":"LinkedIn","url":"https://linkedin.com/in/chandra-prakash-s","icon":"FiLinkedin"},{"name":"GitHub","url":"https://github.com/leviathanaxeislit","icon":"FiGithub"},{"name":"Email","url":"mailto:chandra.1997@hotmail.com","icon":"FiMail"},{"name":"Website","url":"https://www.chandraprakashs.im","icon":"FiGlobe"}],"projects":[{"id":"skills-platform","title":"AI Agent Skills Platform","subtitle":"Multi-Tenant Skill Lifecycle, Bounded Execution & Human-in-the-Loop Approvals","summary":"Production-grade platform for creating, testing, versioning, and executing user-defined AI agent skills with strict permission controls, human-in-the-loop approvals, per-user multi-tenant data isolation, and structured execution logging.","category":"Full-Stack AI Architecture • Multi-Tenant Systems & Agentic Workflows","problem":"Executing user-defined AI skills introduces security risks (unbounded API calls, mutating database actions without approval), tenant data leakage in multi-user environments, and non-reproducible skill executions due to unversioned mutations.","solution":"Architected a Clean Architecture platform with Next.js 16 (App Router) and FastAPI. Engineered a bounded execution engine with step limits and retry capping (max 3), per-user multi-tenant data isolation, human-in-the-loop approval workflows with idempotency protection for mutating tools, immutable versioning (Draft -> Published v1 -> v2 draft bump), Supabase JWT authentication, Gemini LLM sample input generation, and structured logging (structlog).","impactMetrics":["100% per-user data isolation with composite (user_id, name) uniqueness across Supabase Auth & PostgreSQL","Bounded execution engine enforcing max step constraints, tool permission validation, and 3-retry resilience","Human-in-the-Loop approval state machine with idempotency protection for mutating operations","Immutable versioning lifecycle preserving historical execution reproducibility (Draft -> Published v1)","Comprehensive 111-test suite with pytest covering unit, integration, permission, and approval workflows"],"tags":["Next.js 16","FastAPI","Python","TypeScript","SQLAlchemy 2.0","Pydantic v2","Supabase Auth","PostgreSQL","Google Gemini API","structlog","Docker"],"image":"/images/aiskillsmanager.png","featured":true,"liveUrl":"https://ai-skills-manager.vercel.app/"},{"id":"clarityux","title":"ClarityUX","subtitle":"Multi-Provider AI Inference & Visual Attention Analytics","summary":"Production AI inference system and OpenCV computer vision pipeline for predictive heatmaps, attention hotspots, and eye-tracking simulations serving 4,000+ active users across Figma and Chrome.","category":"Company Experience • AI Vision & LLM Orchestration","company":"ClarityUX","role":"Software Engineer","problem":"Production AI workflows often suffer from vendor lock-in, latency bottlenecks, and subjective UI review iterations lacking empirical visual attention data.","solution":"As a Software Engineer at ClarityUX, architected a multi-provider inference system integrating Gemini, Claude, Groq, and OpenAI alongside OpenCV/FastAPI computer vision pipelines for predictive heatmaps, attention hotspots, and eye-path simulations across Figma plugins, Chrome extension, and Next.js platforms. Shipped production Conversational AI systems, Generative AI workflows, and real-time user feedback systems.","impactMetrics":["40% reduction in UX analysis turnaround time","Serving 4,000+ active users globally across Figma & Chrome","Engineered Conversational AI assistants & Generative AI UI synthesis features","Built real-time user feedback system for continuous UX evaluation & iteration","30% reduction in backend bottlenecks with Node.js inference pipeline"],"tags":["Conversational AI","Generative AI","Feedback System","Python","FastAPI","OpenCV","TensorFlow","React","Next.js","LiveKit"],"image":"https://framerusercontent.com/assets/0OqBZwAAiMMh5InAmupzRuYXbQ.png","featured":true,"liveUrl":"https://www.clarityux.com"},{"id":"altruisty","title":"AI Workforce Planning","subtitle":"Dual-Sided Workforce Intelligence & Recommendation Engine","summary":"Dual-sided workforce intelligence platform combining Collaborative Filtering and PyTorch POCs scaled to Gemini Embeddings and Firestore Vector Search for semantic candidate ranking, JD synthesis, and candidate companion tools.","category":"Generative AI, Recommendation Systems & Vector Search","company":"Altruisty","role":"ML & Systems Architect","problem":"High-volume hiring pipelines suffer from mismatched candidate recommendations, manual resume evaluations, and fragmented recruiter outreach, while candidates lack tailored career insights and instant guidance.","solution":"Architected a dual-sided workforce intelligence platform. Built initial POCs utilizing PyTorch, TensorFlow, and Collaborative Filtering models before scaling to production with Gemini Embeddings and Firestore Vector Search. For recruiters, provided semantic resume matching, KB role auto-JD generation, candidate Kanban board, and cold message email dispatch. For candidates, engineered a role recommendation portal providing salary/career path insights, ATS resume builder, cover letter generation, and a companion Telegram bot for mobile assistance.","impactMetrics":["Architected POC models using PyTorch, TensorFlow & Collaborative Filtering before migrating to production Gemini Embeddings","Recruiter Dashboard: Firestore Vector Search for semantic resume ranking, role KB auto-JD generator & interactive Candidate Kanban","Candidate Portal: Personalized role recommendations with salary benchmarks, career path insights & tailored job discovery","AI Career Copilot & Telegram Companion Bot: Auto-generates ATS resumes, cover letters, and provides mobile candidate assistant alerts","Recruiter Outreach: Gemini cold message template synthesis with direct-from-dashboard candidate email dispatch"],"tags":["Gemini API","Gemini Embeddings","Firestore Vector Search","Telegram Bot Companion","PyTorch","TensorFlow","Collaborative Filtering","Python","Next.js","ATS Resume Builder","Recruiter Kanban"],"image":"/images/aiwpt/1.png","gallery":["/images/aiwpt/1.png","/images/aiwpt/2.png","/images/aiwpt/3.png","/images/aiwpt/4.png","/images/aiwpt/5.png","/images/aiwpt/6.png","/images/aiwpt/7.png"],"featured":true,"liveUrl":"https://aiwpt.chandraprakashs.im"},{"id":"sunflower-ai","title":"Sunflower AI","subtitle":"Voice-First Companion & Interactive Memory Garden","summary":"Voice-first conversational assistant pairing Gemini Live API real-time audio streaming with a Three.js 3D avatar and long-term memory engine for reflective journal synthesis.","category":"Voice AI & Interactive Companion","problem":"Most voice assistants feel cold, robotic, and transactional—forgetting conversation context as soon as a session ends.","solution":"Built Sunflower AI, a voice-first companion designed to help users feel heard and supported. It combines a live 3D avatar that lip-syncs in real time with a continuous memory engine that remembers past conversations and grows a visual Memory Garden.","impactMetrics":["Real-time voice streaming using Gemini Live API with instant lip-sync animations","Smart memory engine that recalls personal context and generates reflection journals","Ghibli-inspired 3D room and interactive Memory Garden built with Next.js & Three.js"],"tags":["Next.js 15","React Three Fiber","Gemini Live API","Web Audio API","Firebase","Zustand","Tailwind CSS"],"image":"/images/sunflower.png","featured":true,"liveUrl":"https://sunflowerai.chandraprakashs.im/"},{"id":"android-kernel","title":"Android Open Source & Custom Kernel","subtitle":"AOSP Kernel Development & Low-Level Tuning","summary":"Custom Linux kernel and AOSP distribution for mobile hardware featuring custom CPU schedulers, governor tuning, and low-level rendering optimizations serving 1,000+ active users.","category":"Systems & Low-Level Engineering","problem":"Stock Android ROMs often lack fine-grained kernel scheduling, battery efficiency, and hardware performance tuning for specific mobile chipsets.","solution":"Developed and maintained a custom Linux kernel for Xiaomi Redmi Note 4 (mido) and maintained custom AOSP ROMs for Xiaomi Mi A2, implementing custom CPU schedulers, battery optimizations, and Gaussian blur rendering.","impactMetrics":["Serving 1,000+ active users across Android custom ROM ecosystem","Implemented custom Linux kernel schedulers & governor battery tuning","Pioneered early Gaussian blur UI rendering before native Android support"],"tags":["C","Linux Kernel","Android AOSP","Git","System Performance"],"image":"https://images.unsplash.com/photo-1607252650355-f7fd0460ccdb?q=80&w=1000&auto=format&fit=crop","featured":false,"liveUrl":"https://github.com/leviathanaxeislit"}],"articles":[{"id":"multi-provider-llm-orchestration","title":"Architecting Multi-Provider LLM Orchestration in Production","summary":"How to build resilient AI inference pipelines integrating Gemini, Claude, Groq, and OpenAI with automated fallbacks, cost optimization, and unified streaming interfaces.","content":"In production AI engineering, single-provider dependency is a critical vulnerability. Rate limits, regional outages, API degradation, and unexpected latency spikes can halt core user workflows. By implementing dynamic model routing, automated circuit breaking, fallback chains across Groq, Gemini, Claude, and OpenAI, and low-latency stream normalization, systems achieve 99.99% operational inference reliability while cutting costs by up to 35%.","category":"AI Architecture","date":"May 2026","readTime":"6 min read","image":"/images/illustrations/llm_orchestration.png","featured":true,"tags":["LLM Orchestration","Python","FastAPI","Circuit Breakers","Streaming","Groq","Gemini","Claude"],"takeaways":["Single-provider API dependencies create single points of failure in AI production apps.","Tiered fallback matrix (Groq -> Gemini -> Claude -> OpenAI) delivers 99.99% uptime.","Token stream normalization decouples frontend rendering from upstream vendor SDK churn.","Cost-aware model routing reduces total token expenditure by 35% without sacrificing accuracy."],"sections":[{"title":"The Fragility of Single-Provider Infrastructure","content":"Relying on a single AI provider in production introduces severe operational risk. HTTP 429 rate limit exceptions, unannounced API deprecations, latency spikes during peak hours, and regional outages disrupt user sessions. When an app relies on LLM inference for visual attention scoring, UI component generation, or conversational assistants, downtime directly damages user trust. To build production-grade infrastructure, model endpoints must be treated as interchangeable worker nodes within a managed routing matrix.","callout":"Architecture Principle: Treat individual model endpoints as disposable worker instances. High availability requires decoupling business logic from vendor SDKs."},{"title":"Designing the Multi-Provider Fallback Matrix","content":"We established a multi-tiered routing topology based on latency profiles, context windows, and cost efficiency. Ultra-fast providers (Groq running Mixtral and Llama 3) serve low-latency token extraction and initial user feedback loops. High-reasoning workflows route to Gemini 1.5 Pro and Claude 3.5 Sonnet. If an upstream provider fails health checks or exceeds latency thresholds, the circuit breaker automatically routes requests down the fallback chain."},{"title":"Token Stream Normalization & Uniform Event Dispatch","content":"A major challenge in multi-provider architectures is payload variance across vendor streaming interfaces. OpenAI uses Server-Sent Events (SSE) with nested delta objects, Anthropic uses typed event frames (content_block_delta), and Google Generative AI streams raw proto-like response chunks. We built a lightweight stream normalization layer in FastAPI that intercepts raw vendor chunks and emits a uniform event protocol to the client containing delta text, finish reasons, and token usage metrics."},{"title":"Production Uptime & Performance Results","content":"In production deployment across thousands of active requests, the multi-provider orchestrator achieved 99.99% inference uptime. Average time-to-first-token (TTFT) dropped by 40% when routing brief queries to Groq, and total token expenditure decreased by 35% through intelligent cost-threshold routing."}]},{"id":"computer-vision-attention-heatmaps","title":"Predictive Visual Attention with OpenCV & FastAPI","summary":"Using computer vision pipelines to simulate human eye-tracking paths and predict UI hotspots before user testing.","content":"Traditional eye-tracking studies are expensive, slow, and unscalable during rapid product iteration. By leveraging OpenCV feature detection, spectral visual saliency algorithms, and color space mathematical transformations exposed via asynchronous FastAPI endpoints, product teams receive instant visual attention heatmaps and CTA focus scores directly inside Figma and browser extensions in under 50ms.","category":"Computer Vision","date":"Feb 2026","readTime":"5 min read","image":"/images/illustrations/visual_attention.png","featured":false,"tags":["OpenCV","Computer Vision","FastAPI","Python","Predictive Heatmaps","Visual Attention","NumPy"],"takeaways":["Replaced manual 2-week eye-tracking studies with instant algorithmic vision analysis.","Spectral Residual analysis isolates high-contrast visual entry points across UI layouts.","NumPy array operations and OpenCV C++ bindings process 1080p UI screenshots in under 45ms.","Empowered 4,000+ Figma & Chrome users to optimize visual hierarchy prior to deployment."],"sections":[{"title":"Eliminating Empirical Eye-Tracking Bottlenecks","content":"Design review iterations frequently stall due to subjective opinions on visual hierarchy. Physical eye-tracking labs require expensive hardware, participant recruitment, and days of data collection. To solve this, we engineered an algorithmic vision pipeline that predicts visual attention hotspots automatically by analyzing luminance, color opponency, and structural edge density in UI screenshots.","callout":"Product Vision: Provide quantitative, empirical visual feedback at the moment of design creation rather than post-launch user analytics."},{"title":"The Vision Algorithm Pipeline","content":"The image processing pipeline operates in three stages: Image Decomposition, Spectral Residual Saliency Computation, and Gradient Heatmap Rendering. Input screenshots are transformed into grayscale and converted to the frequency domain using Discrete Fourier Transform (DFT). The spectral residual represents the visual surprise—regions that diverge from the natural image frequency spectrum—which corresponds strongly to human visual attention focus."},{"title":"Low-Latency Service Architecture","content":"Exposing heavy computer vision models to real-time design plugins requires tight performance constraints. By utilizing in-memory NumPy byte buffers and asynchronous FastAPI worker threads, the service processes 1080p UI frames, generates heatmap overlays, and returns encoded PNG payloads in under 45ms, enabling real-time preview as designers drag elements across their Figma canvas."}]},{"id":"kernel-optimizations-aosp","title":"Lessons from Linux Kernel & Custom ROM Development","summary":"What operating system kernel scheduling and memory management teach us about modern backend engineering.","content":"Writing custom CPU schedulers, memory governor tuning, and low-level rendering optimizations for Android custom ROMs instills deep respect for hardware efficiency. Operating at the metal level changes how you view backend architecture: thread synchronization, cache line alignment, memory allocation overhead, and non-blocking I/O become first-class concerns.","category":"Systems Engineering","date":"Jan 2026","readTime":"7 min read","image":"/images/illustrations/kernel_optimizations.png","featured":false,"tags":["Linux Kernel","C","Android AOSP","Low-Level Tuning","Systems Architecture","Memory Management"],"takeaways":["Hardware constraints build disciplined software engineers who reject unnecessary abstraction.","Energy Aware Scheduling (EAS) governor tuning achieved 60fps UI smoothness while lowering power draw by 18%.","zRAM memory compression principles directly translate to Redis and database buffer pool optimization.","The best system engineering is invisible—users only notice that the interface feels instantaneous."],"sections":[{"title":"Confronting the Hardware Reality","content":"In high-level application development, virtual machines and garbage collectors mask resource management. In mobile Linux kernel development, every unaligned memory access, lock contention, or unnecessary CPU frequency spike directly drains physical battery life and drops UI frames. Tuning low-level Android distributions for thousands of active users cultivated a relentless focus on efficiency and zero-overhead design.","callout":"Systems Philosophy: Software efficiency is not an afterthought or a micro-optimization; it is a fundamental design requirement."},{"title":"Custom CPU Governor & Frequency Hysteresis Tuning","content":"Default Linux governors often react poorly to bursty user touch events, resulting in stuttered frame rendering or aggressive battery drain. By modifying Energy Aware Scheduling (EAS) energy models in C and tuning frequency ramp-up hysteresis, we ensured the CPU frequency scaled instantly upon touch input while holding stable during micro-pauses."},{"title":"Kernel Principles Applied to Backend Engineering","content":"The low-level mechanics of kernel development directly shape how we architect modern distributed microservices:\n\n1. Interrupt Handling & Event Loops: Kernel IRQ top-half/bottom-half separation mirrors non-blocking Node.js/FastAPI event-driven architecture.\n2. Memory Reclamation (zRAM): Compressing RAM pages before swapping to disk informs zero-copy data serialization and Redis memory compaction.\n3. Governor Hysteresis: CPU scaling cooldown logic directly applies to cloud auto-scaling metrics to prevent server thrashing during traffic spikes."}]},{"id":"realtime-voice-ai-streaming","title":"Building Ultra-Low Latency Voice AI with Gemini Live API & Web Audio","summary":"Architecting full-duplex real-time audio streaming, WebSocket PCM buffer management, and Three.js 3D avatar lip-sync synchronization.","content":"Conventional voice assistants rely on multi-stage pipelines: Speech-to-Text -> Text LLM -> Text-to-Speech synthesis. This turn-based architecture introduces 2-4 seconds of latency, creating a jarring, mechanical experience. By implementing full-duplex WebSocket audio streaming powered by Gemini Live API and synchronized with Web Audio API viseme processing, we achieved sub-300ms natural conversational audio flow with a 3D visual avatar.","category":"Voice AI & Real-time Systems","date":"Jul 2026","readTime":"6 min read","image":"/images/illustrations/voice_ai_streaming.png","featured":false,"tags":["Voice AI","Gemini Live API","Web Audio API","WebSockets","Three.js","React"],"takeaways":["Full-duplex WebSocket audio streaming reduces conversational turn latency under 300ms.","Web Audio Worklet nodes handle zero-pop PCM 24kHz playback buffer queuing.","Real-time audio frequency analysis drives Three.js morph target lip-sync visemes.","Empathy-first conversational AI requires low-latency streaming and persistent user memory."],"sections":[{"title":"The Latency Wall in Voice Conversational Systems","content":"Human conversation relies on subtle timing cues—overlaps, brief pauses, and back-channel acknowledgments happen in under 300 milliseconds. Traditional sequential voice stacks (whisper STT -> GPT text completion -> ElevenLabs TTS) break down because latency compounds at every layer. Full-duplex streaming over WebSocket transport eliminates discrete audio boundaries, allowing user input and AI audio output to flow continuously.","callout":"Design Goal: Fluid conversational timing where users can interrupt mid-sentence and receive immediate acoustic response without robotic delays."},{"title":"PCM Buffer Queuing with Web Audio API","content":"Receiving streaming base64 PCM 24kHz audio chunks over WebSockets requires precise client-side timing to prevent audio crackles or starvation gaps. We built a custom AudioStreamHandler using the Web Audio API that queues incoming PCM float32 buffers and schedules contiguous playback timestamps."},{"title":"Real-time Avatar Lip-Sync & Memory Garden Integration","content":"As audio streams into the speakers, an AnalyserNode computes fast Fourier transforms (FFT) to extract instantaneous amplitude and frequency bands. These values map to 3D facial morph targets (visemes) on a Three.js Ghibli-inspired avatar in real time. Combined with a persistent Firebase memory engine, the assistant remembers past conversations, creating a deeply human conversational companion."}]}]}