Building AI products, agentic systems, and workflow intelligence tools
Sydney, Australia
Portfolio · Email · LinkedIn · GitHub
I build 0→1 AI products that turn messy real-world inputs and high-friction decisions into systems people can understand, operate, and trust.
Most recently, I built Orchestra through Arrayah and SH1P Australia in Sydney. Orchestra is a citation-grounded Product Memory for software teams, designed to turn meetings, documents, conversations, and code into searchable project knowledge. I took it from 100+ customer interviews to live beta with five pilot customers, owning product strategy, system architecture, evaluation, and engineering delivery.
I’m currently pursuing a Master of Data Science and Innovation at UTS, and I’m especially interested in:
- agentic workflows and tool use
- RAG, hybrid retrieval, grounding, and citations
- LLM evaluation and observability
- full-stack AI products
- workflow automation and AI governance
A private product built through Arrayah and SH1P Australia.
What it is
- A citation-grounded Product Memory for software teams
- Turns meetings, documents, conversations, and code into searchable project context
- Uses hybrid retrieval and reranking across structured and unstructured sources
- Includes a 148-case evaluation suite for relevance, grounding, and regression detection
- Connects through an MCP server and 17 read-only integrations across Slack, GitHub, and Google Drive
- Progressed from customer discovery to live beta with five pilot customers
- Start with the user, decision, existing workaround, and smallest valuable workflow—not a model looking for a use case
- Separate model proposals from software authority, permissions, approvals, and consequential actions
- Evaluate retrieval, grounding, task quality, safety, latency, and cost as different product risks
- Build deterministic local or mock paths so the product, tests, and demos remain reproducible without paid providers
- Treat observability, failure states, auditability, and human escalation as product surfaces rather than backend afterthoughts
Sydney, Australia | Mar 2026 – Jun 2026
- Took Orchestra from 100+ customer interviews to live beta with five pilot customers
- Owned product discovery, system architecture, retrieval, evaluation, backend delivery, and integration strategy
- Built hybrid retrieval and reranking, a 148-case evaluation suite, an MCP server, and 17 read-only integrations
Sydney, Australia | Feb 2026 – Mar 2026
- Built a privacy-first AI resume-screening system using section-aware scoring, dynamic weighting, and semantic clustering
- Designed an explainable LLM refinement layer while keeping document processing local and every score auditable
Remote | Dec 2024 – Jan 2025
- Engineered NLP and speech-recognition pipelines for a real-time voice-to-voice translation system
- Reached 97% language-detection accuracy and improved contextual translation quality
Bengaluru, India | May 2024 – Oct 2024
- Built a real-time CNN-based terrain-detection system combining camera and sensor data for automotive safety
- Deployed a RAG assistant over internal engineering documentation, adopted by 3,000+ employees
- Automated airbag-testing documentation in Python, reducing manual effort by roughly 95%
Risk-aware model and agent routing for Codex
- Solves the product problem of choosing the right model, specialist, and review depth without making developers reason about the entire agent stack
- Analyses task and repository context, evaluates complexity and risk, and selects an appropriate execution lane
- Routes work across GPT-5.6 model lanes and 172 bundled specialist agents while keeping the lane decision separate from implementation
- Produces auditable route cards, bounded execution paths, and verified handoffs backed by tests and evidence
Governed natural-language analytics workspace
- Gives business teams a faster path to answers without granting generated SQL direct authority over sensitive data
- Turns business questions into reviewable SQL, visualisations, churn analysis, forecasts, and executive reports
- Treats generated SQL as an untrusted proposal and constrains it through table and column policy, immutable approval binding, RBAC, audit trails, and execution budgets
- Built with FastAPI, Next.js, DuckDB, SQLGlot, PostgreSQL, and MLflow
AI workflow control plane for high-stakes operations
- Makes operational automation inspectable and recoverable instead of hiding consequential actions inside model-generated prose
- Compiles operational evidence into typed workflows that people can inspect, approve, execute, recover, and replay
- Keeps models in a proposal role while deterministic policy controls approvals, idempotent effects, postcondition verification, and evidence exports
- Includes a live product, deterministic evaluation path, recovery workflow, and hash-chained execution traces
Role-aware RAG, human review, and AI governance platform
- Helps knowledge workers reach concise answers without treating a fluent model response as the source of truth
- Turns approved documents into grounded answers with citations, retrieved evidence, confidence scoring, and agent traces
- Routes low-confidence answers to human review and records evaluation, audit, usage, latency, and operational signals
- Built with FastAPI, Next.js, PostgreSQL, pgvector, Redis, Docker, JWT/RBAC, and CI
- Manages evaluation datasets, agent runs, human reviews, quality gates, and observability in one workspace
- Makes changing AI behaviour inspectable through versioned evidence rather than isolated prompt experiments
- Orchestrates planning, patching, testing, repair loops, and approval for software-engineering agents
- Uses sandboxed commands, secret scanning, deterministic evaluations, and bounded execution controls
- Converts alerts and evidence into triage hypotheses, tool traces, approval-gated actions, and incident reports
- Keeps operational actions mock-first and reviewable through RBAC, audit, evaluation, and observability
- Supports voice-led customer service with contextual tools, supervisor handoff, and audited call workflows
- Combines a real-time product surface with role-aware operations, evaluation, and optional model providers
- Creates district-level agricultural intelligence across climate, soil, water, crop, and policy risk
- Combines geospatial exploration, crop recommendations, economic comparisons, and policy simulation
- Voice-first AI travel companion for backpackers exploring Australia
- Combines weather-aware recommendations, maps, budgets, itinerary planning, travel utilities, and emergency references
- Available as a live product
- Turns a student’s PDFs and slides into evidence-linked quizzes, explanations, and revision plans
- Validates citations against retrieved context and preserves a deterministic, no-key study workflow
- Intent-aware data-science copilot for profiling, cleaning, visualisation, baseline modelling, and export
- Keeps transformations reviewable through explicit cleaning logs and local-first data handling
- Local-first resume screening using section-aware evidence rather than naive keyword overlap
- Combines deterministic scoring, semantic clustering, matched and missing requirements, and optional LLM explanations
- Turns campaign briefs into full-funnel copy, video concepts, validation, feedback loops, and a Notion-backed review queue
- Keeps brand, legal, and publishing decisions with a human reviewer
- Selected for Arrayah Accelerator Chapters 3 & 4 and SH1P Australia Cohort 1 to build Orchestra
- Submitted a technical paper on the AI-Enhanced Terrain-Adaptive Vehicle Control System for the SAEINDIA International Mobility Conference 2024
- Showcased the terrain-adaptive vehicle system at Continental Innovation Day
- Published "Real-Time Biometrics-Based Smart EVM with FPGA Implementation" in the International Journal of Scientific Research and Engineering Trends
- Built AgriSmart as part of the Mistral Hackathon in Sydney
Languages
Python, TypeScript, SQL
AI Systems
Agent orchestration, RAG, hybrid retrieval, reranking, embeddings, vector search, tool calling, MCP, structured outputs, prompt engineering, grounding, citation verification, LLM evaluation
Full-Stack & Data
FastAPI, Pydantic, Node.js, React, Next.js, PostgreSQL, pgvector, Redis, DuckDB, SQLGlot, MLflow, Pandas, NumPy, scikit-learn
Platform & Reliability
Docker, CI/CD, GitHub Actions, JWT/RBAC, audit logs, observability, sandboxing, security scanning, OpenAI, Anthropic, Groq
Master of Data Science and Innovation
2025 – 2027
Bachelor of Engineering in Electronics and Communication
2020 – 2024
- Portfolio: karthikrameshportfolio.vercel.app
- Email: karthikramesh2012@gmail.com
- LinkedIn: karthik-ramesh-2b52ab328
- GitHub: KarthikRamesh9149
I’m interested in teams turning difficult customer workflows into dependable AI products—especially where product judgment, hands-on implementation, and responsible automation need to work together.