SHRASIT Solutions
Services · AI & ML · Applied AI

Retrieval-grounded copilots, tool-using agents, document intelligence and forecasting, built for enterprises that must defend the output. Wired into the SAP, Maximo and HIS estates we already implement, with the evaluation, guardrails and MLOps to run them safely in production.

Eval pipeline · every change

Nothing ships ungraded.

Stage 01

hybrid index

Retrieve

Stage 02

grounded only

Generate

Stage 03

golden set

Evaluate

Stage 04

human review

Gate

Regression gateblocks on drop
Scoping
Acceptance criteria are written before training begins, not after the demo.
Grounding
Answers cite your documents and records, or the system declines to answer.
Oversight
A human approves anything that moves money, medication or a legal record.
Operations
Drift monitoring and a rollback path from day one.

Honest scoping

When AI is the answer, and when it isn’t.

Many AI pilots that never reach production were the wrong tool from the start. We have this conversation before the statement of work, not after the budget is spent.

AI wins

  • The input is unstructured — documents, tickets, notes, images — and rules never covered the long tail.
  • An expert already makes the judgement consistently, with a written record of it.
  • Volume is high, so a few points of accuracy convert into real hours or cost.
  • The output can be checked — by a person, a downstream system or a later event.
  • Being approximately right now beats being exactly right next week.
  • Knowledge is scattered across systems, and retrieval can span them without a migration.

Plain software wins

  • The logic is a policy you can write down — a decision table is cheaper and auditable.
  • The task must be exactly right every time, with no reviewer before the consequence.
  • There are no labelled examples, no expert to create them, and no way to judge an answer.
  • A regulator requires a deterministic, reproducible explanation for every decision.
  • The data is too thin, stale or biased, and fixing it is the real project.
  • Volume is low enough that the evaluation harness would cost more than the manual work.

Where it ships

Six patterns, in the industries we already serve.

Each one sits on a system of record we implement and support, which is what makes a copilot change a real workflow rather than just a demo.

  • Clinician reviewing patient notes on a tablet at a hospital ward stationHealthcare

    Clinical documentation & coding

    Draft discharge summaries and suggest ICD codes from the notes already in the HIS. The clinician reviews and signs; the model never writes to the record unattended, and every suggestion is traceable to its source line.

    Approach

    RAGStructured extractionHuman in the loop
    Industry context
  • Banking operations floor with analysts working across multiple screensBanking & FSI

    Document intelligence for onboarding & credit

    Read KYC packs, trade documents and financial statements at intake. Fields are extracted with confidence scores; low-confidence items route to an analyst. Cuts the re-keying that slows onboarding without moving the approval decision to a model.

    Approach

    Document AIExtractionConfidence routing
    Industry context
  • Government ministry building with reflective glass facadeGovernment

    Citizen-service assistants

    One assistant over the circulars, forms and eligibility rules spread across ministry portals. Answers cite the source document and version, in Arabic and English, with escalation to a human on anything contested.

    Approach

    RAGCitationsBilingual retrieval
    Industry context
  • Refinery pipework and rotating equipment under maintenance inspectionOil & Gas

    Predictive maintenance on Maximo

    Failure prediction on rotating equipment from sensor history and the work-order record in IBM Maximo. Forecasts arrive as scheduled work orders, not a dashboard — the EAM integration is the point.

    Approach

    Time-series MLAnomaly detectionEAM integration
    Industry context
  • Robotic arms inspecting components on an automated assembly lineManufacturing

    Visual inspection & demand forecasting

    Edge-run defect detection on the line, plus demand and capacity forecasts with confidence intervals, so planners see a range rather than a single false-precision number.

    Approach

    Computer visionForecastingEdge inference
    Industry context
  • Support agents wearing headsets at a customer service operations deskRetail & Telecom

    Support copilots on your own history

    Agent-assist grounded in resolved tickets, product data and policy. It drafts a reply, cites its source, and hands over when confidence drops. Deflection is measured on a held-out set before it is claimed.

    Approach

    RAGAgent assistEval-gated rollout
    Industry context

Approaches

Four approaches, chosen by the problem.

A gradient-boosted model on structured data still beats a language model at most enterprise prediction tasks, at a fraction of the cost. We choose on evidence and build so the model underneath can be replaced.

Retrieval-augmented generationApproach

Your documents, answered with citations.

The default for knowledge work. Chunking and embeddings tuned to your corpus, hybrid keyword-plus-vector retrieval, re-ranking, and generation constrained to retrieved context. Access control is enforced at retrieval time, so the model never surfaces a document the user could not open.

Citations by defaultNo retraining to add knowledgeRow-level access controlAuditable answer provenance

Typical use case · Policy assistants, support copilots, ministry knowledge bases.

Tool-using agentsApproach

Models that act within limits you define.

An agent plans, calls only the tools you allow, and stops. Each tool is a typed function with its own permissions and audit log; write operations sit behind human approval. Scope is kept narrow by design.

Typed, permissioned toolsApproval gates on writesFull call audit trailBounded retries and budgets

Typical use case · Ticket triage, reconciliation, multi-system lookups.

Classical ML & forecastingApproach

Structured data, explainable predictions.

Gradient boosting, time-series models and survival analysis for demand, capacity, churn, credit and failure prediction. Faster, cheaper and easier to explain to a regulator than an LLM, and our first choice when the input is already structured.

Explainable feature importanceCheap to retrainConfidence intervalsRuns on modest hardware

Typical use case · Demand planning, predictive maintenance, credit risk.

Computer vision & document AIApproach

Pixels and PDFs into structured records.

Layout-aware document parsing, OCR for Arabic and English, table extraction, and inspection models that run at the edge. Output is structured fields with per-field confidence, so downstream systems can route uncertainty rather than swallow it.

Layout-aware parsingArabic + English OCRPer-field confidenceEdge and on-prem deployment

Typical use case · Invoice and KYC intake, defect detection, meter reading.

Global network and data infrastructure visualization at night

Why pilots stall

Retrieval over a messy document estate, permissions that survive an audit, an evaluation harness the business trusts, and an integration into the system of record. That is where the months go, and it is the part a demo never shows.

Delivery method

Five steps from framing to running in production.

The evaluation harness is built before the system it grades. No rollout without a measured baseline, and no launch without a rollback someone has actually rehearsed.

  1. 01

    Frame

    A two-week diagnostic to name the decision the model should change and who acts on it. If a decision table or report would do the job, we say so.

    Tooling

    • Use-case canvas
    • Baseline measurement
    • Feasibility spike
  2. 02

    Ground

    Data access, lineage, access controls, PII handling and retrieval design — usually the longest phase, and the one that decides whether the rest works.

    Tooling

    • Data contracts
    • Vector index design
    • PII redaction
  3. 03

    Evaluate

    Build the evaluation harness before the system: held-out sets, expert-labelled goldens, adversarial cases and an accuracy bar agreed in advance.

    Tooling

    • Golden datasets
    • LLM-as-judge + human review
    • Regression suite
  4. 04

    Ship

    Integrate into the system of record — SAP, Maximo, HIS, ServiceNow — behind guardrails and approval gates. Shadow mode first, then a measured rollout with a kill switch.

    Tooling

    • Shadow deployment
    • Approval workflow
    • Feature flags
  5. 05

    Operate

    Drift and quality monitoring, per-request cost tracking, scheduled re-evaluation, and a rehearsed rollback to the previous model version.

    Tooling

    • MLflow registry
    • Drift monitors
    • Cost dashboards

Guardrails & evaluation

The work that keeps AI in production.

A model that is right nine times in ten is a liability without the tenth case handled. Evaluation, grounding and human review are engineering work, not a closing slide.

Grounding & citations

Answers are constrained to retrieved context and cite it. When nothing relevant is found, the system says so.

Evaluation harness

Golden sets, adversarial prompts and regression tests run on every prompt or model change, like a CI suite for code.

Human in the loop

Anything touching money, medication, employment or a legal record gets an approval step. The model drafts; a person signs.

PII & access control

Redaction before inference, permissions enforced at retrieval, and per-tenant isolation, so a chat box never exposes a document the user could not open directly.

Prompt-injection defence

Untrusted content is never treated as instructions. Tools stay least-privilege, and outputs are validated against a schema before reaching another system.

Drift & cost monitoring

Quality tracked against a live sample and spend tracked per request and team, with alerts before a bill or error rate becomes a problem.

Evaluation trail

Six gates, every program
  1. Gate 01

    Frame

    Accuracy bar, refusal policy and the human-review boundary agreed in writing before any build starts.

  2. Gate 02

    Baseline

    Measure how the process performs today. Without it, no later number means anything.

  3. Gate 03

    Offline

    Golden datasets and adversarial cases. Prompt and model changes blocked at the gate on regression.

  4. Gate 04

    Human

    Domain experts score a sampled set blind. Disagreement between reviewers is tracked, not averaged away.

  5. Gate 05

    Shadow

    Runs live against real traffic with output withheld. Compared against the human decision, no user impact.

  6. Gate 06

    Operate

    Scheduled re-evaluation, drift alerts and a rehearsed rollback to the last known-good version.

Tech stack

The tools we ship and operate with.

Python and PyTorch where models are trained, hosted and open-weight LLMs behind one interchangeable interface, MLflow for the registry, and the same Snowflake and Databricks foundations our data teams already run.

Python
PyTorch
scikit-learn
XGBoost
Hugging Face
Transformers
OpenAI
Anthropic
Azure OpenAI
Amazon Bedrock
Vertex AI
LangChain
LlamaIndex
vLLM
Ollama
pgvector
Pinecone
Weaviate
FAISS
MLflow
Weights & Biases
Ray
Airflow
dbt
Snowflake
Databricks
Delta Lake
FastAPI
Docker
Kubernetes
Power BI

FAQ

Frequently asked questions.

  • How do you stop a copilot from confidently making things up?

    Generation is constrained to retrieved context, so the model answers from your documents, not its training data. Every answer carries a citation you can open and check, and the system is allowed to refuse when nothing relevant is found. Adversarial cases designed to bait a hallucination run on every prompt change.

  • Hosted models or self-hosted open-weight models?

    Whatever your data residency and regulator allow, decided at the framing stage. Hosted APIs (OpenAI, Anthropic, Azure OpenAI, Bedrock) reach quality faster and cost less at moderate volume; open-weight models on your own infrastructure (vLLM, Ollama) suit data that cannot leave the country or building. The retrieval and evaluation layers are built so the model underneath can be swapped.

  • Our data is scattered and half of it is in PDFs. Is that a blocker?

    No — that is the normal starting point. Layout-aware parsing, OCR for Arabic and English, and a hybrid index handle messy document estates without a migration project. What does block a program is data nobody owns, or a corpus too contradictory to answer defensibly, which we surface in the framing phase.

  • When do you tell a client not to use AI?

    When the rule is writable — a decision table is cheaper, auditable and does not drift. We also decline when the output cannot be evaluated, when a regulator needs a deterministic explanation for every decision, or when the real project is fixing the data foundation.

  • Can this integrate with our SAP, Maximo or HIS estate?

    Yes, and that is usually where the value is — a forecast that lands as a scheduled work order in Maximo changes operations; the same forecast on a dashboard does not. We already implement and support these systems, so the adapters, writers, reconciliation and permission mapping are familiar rather than a discovery exercise.