← All insights

Insights

Why Most Malaysian AI Projects Fail Before They Start: A Data Quality Checklist

Most Malaysian AI projects do not fail because the model is weak. They fail earlier — when copilots, RAG, or predictive models are pointed at data that is incomplete, inconsistently defined, or quietly full of personal data no one mapped.

The board sees a polished demo on a clean sample. Production sees nulls, duplicate customers, three versions of “active”, and a spreadsheet that only one person understands. Trust collapses. Budget freezes. The narrative becomes “AI doesn’t work here.”

Usually, AI was never the first problem. Data readiness was.

The Pattern We See

A typical sequence:

  1. Leadership wants GenAI, a chatbot, or “ask your data.”
  2. A vendor wires a model to a warehouse, SharePoint dump, or CRM export.
  3. Answers look impressive for two weeks.
  4. Business users find hallucinations, wrong numbers, or unsafe personal data in the wrong place.
  5. The project is paused “for review” — and rarely recovers cleanly.

Cross-sell and upsell models fail the same way. If customer keys are broken, product hierarchies disagree, or CRM and warehouse definitions diverge, the model learns noise. Tableau and Power BI then publish that noise with nicer charts.

Fixing this after go-live is expensive. Checking it before kickoff is not.

A Practical Data Quality Checklist Before AI

Use this as a go / no-go gate. If you cannot answer these with evidence, you are not ready to train or ground an AI system on that domain.

1. Completeness and missing data

  • Which critical fields are null or blank in the tables AI will use?
  • Is “missing” random, or concentrated in certain branches, products, or years?
  • Can the business survive answers that skip those records — or will users assume silence means “zero”?

2. Accuracy and uniqueness

  • Do customer, account, or product IDs uniquely identify one real-world entity?
  • How many duplicates, orphans, and broken foreign keys exist in the join paths RAG or models need?
  • Are free-text fields (remarks, notes, complaint descriptions) usable — or mostly noise?

3. Timeliness and freshness

  • How old is the data the AI will see?
  • Which feeds fail silently overnight?
  • If a dashboard or agent says “current,” what is the actual last successful load?

4. Consistency across systems

  • Does “customer” mean the same thing in CRM, billing, and the warehouse?
  • Do finance and sales share one definition of revenue, active, or churn?
  • Where do Excel “shadow systems” still win over the official source?

5. Fitness for AI / RAG specifically

  • Is there enough high-quality grounding content for the questions users will ask?
  • Are documents versioned, owned, and current — or a SharePoint graveyard?
  • Have you listed known blind spots (domains with no reliable source)?

6. Definitions and ownership

  • Who owns each critical entity and metric?
  • Is there a written definition — or tribal knowledge?
  • When IT and business disagree, whose version will the AI learn?

7. PII and PDPA exposure

  • Where does personal data appear outside expected systems?
  • Who can access it once it is embedded, indexed, or logged by an AI tool?
  • Are retention and purpose limitation still true after you expand AI access?

8. Lineage and integration trust

  • Can you draw source → interface → ETL → mart → consumer for the critical path?
  • Which integrations break first under load?
  • If an answer is wrong, can you trace which feed caused it?

How StringRay Turns the Checklist into Evidence

StringRay is Oxydata’s data quality audit offering before AI. It is designed for teams that want findings with evidence, not another opinion workshop.

StringRay typically runs six audits:

  1. Data Quality Audit — missing data, completeness, accuracy, uniqueness, timeliness, consistency.
  2. AI / RAG Data Fitness — coverage, freshness, grounding docs, and blind spots for copilots and models.
  3. Semantic & Definition Alignment — ownership, definitions, and stewardship gaps that break trust.
  4. PII & PDPA Exposure Review — personal data in unexpected places, access risks, handling gaps.
  5. What to Cleanse — a prioritised backlog (duplicates, broken keys, orphans, free-text mess) with effort and owners.
  6. Data Lineage & Integration Traceability — source systems, interfaces, ETL/feeds, and provenance.

The output is a scored findings pack and a remediation roadmap — so AI, enterprise data / datamart, and AI consulting work starts on foundations you can defend.

What “Good Enough for AI” Looks Like

You do not need perfect data everywhere. You need trusted enough data in the domains you are about to automate.

A healthy kickoff usually means:

  • Critical tables profiled with known null rates and duplicate rates
  • Agreed definitions for the entities the AI will talk about
  • A written list of questions the system should not answer yet
  • PDPA risks identified before embeddings and logs multiply exposure
  • A sequenced cleansing backlog with named owners
  • Clear lineage for the feeds that will ground answers or train models

If those are missing, buying a larger model will not save the project.

A Sensible Sequence for Malaysian Teams

  1. Pick one domain — e.g. customer 360 for cross-sell, HR knowledge for a copilot, or hospital KPIs for dashboards.
  2. Run a focused audit — StringRay or an equivalent evidence-based review.
  3. Cleanse what blocks the use case — not everything in the enterprise.
  4. Then build RAG, agents, or predictive models — and connect them to reporting foundations that stay maintainable.

That order feels slower in week one. It is faster than explaining a failed AI programme in month six.

Conclusion

Malaysian enterprises do not lack access to models. They lack trusted inputs. Cross-sell engines, RAG assistants, and board dashboards all inherit whatever quality — or mess — sits underneath.

Before the next GenAI kickoff, run the checklist. If the answers are thin, start with data readiness. That is exactly what StringRay — Data Quality Audit is for.

Oxydata Software helps Malaysian organisations audit data quality before AI, build warehouses and datamarts on cleaner foundations, and deliver GenAI that stakeholders can trust. Talk to us about a StringRay audit or explore StringRay.