AI Services

Machine learning models trained on your data and held to a baseline

The question is never whether a model can be built. It is whether the model beats the rule of thumb your operations manager already uses, by enough to change a decision.

What is machine learning?

Machine Learning is the practice of training models on your historical records to forecast, score or classify: demand by week, which leads will close, which invoices will run late. It suits Australian organisations with a few years of reasonably clean data and a decision repeated often enough that a small accuracy gain is worth real money.

Get a fixed written quote
Typical timeline
12 to 24 weeks
What drives cost
Data condition first, then the number of models, the integration work to put predictions where decisions are made.
Best for
Frequent, repeated decisions with several years of history behind them
You own
The training pipeline, the model artefacts and the feature definitions
Built with
Forecasting, classification, scoring, MLOps and monitoring

Your handover

Machine learning and language models are not the same tool

Almost everything called AI in the last two years has meant language models. Machine learning in the sense on this page is older and, for numerical problems, considerably better. If you want to know how many pallets to hold in Perth in March, a model trained on your own shipment history will beat any general purpose language model comprehensively, run for a fraction of the cost, and give you a number you can trace back to the inputs that produced it.

  1. 01Data assessment with quality issues quantified
  2. 02Documented baseline the model must beat
  3. 03Feasibility report with a written go or stop recommendation
  4. 04Trained model with validation on unseen time periods
  5. 05Feature definitions documented for reproducibility
  6. 06Scoring pipeline integrated where the decision is made
  • Prediction logging with outcome comparison
  • Drift monitoring, alerting and a retraining schedule
  • Handover documentation and rollback procedure
The two also fail differently, which matters for how you supervise them

The two also fail differently, which matters for how you supervise them. A language model fails by inventing something fluent. A trained predictive model fails by being confidently wrong in a measurable way, which sounds worse and is actually far easier to manage, because you can quantify the error, watch it over time and set thresholds at which a human takes over. Most businesses that ask us about AI for a numbers problem want this, not a chatbot, and the distinction usually saves them both money and disappointment.

Start with the baseline you have to beat

Before any modelling, we establish what the current approach achieves. For forecasting the baseline is often naive: last year's number for the same week, or a rolling average. For lead scoring it is whatever heuristic your sales manager applies, which is frequently better than people expect because it encodes years of pattern recognition. Then the model has to beat that measurably on data it has never seen, and if it does not, we tell you and you keep your money.

More on start with the baseline you have to beat

That step ends more projects than any other, and we think it should. A model that improves forecast accuracy by two percent may be a triumph in a business where inventory is the dominant cost and irrelevant in one where it is not. The right conversation is about the decision downstream: what would you do differently if the number were better, and what is that difference worth per year. If nobody can answer that, the model will be built, admired, and ignored.

How the engagement runs

How a machine learning project runs

We work in phases with a genuine decision point after the feasibility work, so you can stop after a small spend rather than committing to a full build on optimism. Roughly a third of the projects we assess do not proceed past that gate, usually because the data cannot support the question being asked.

  1. 01Decision framingWhich decision changes, who makes it, and what better accuracy is worth in dollars per year
  2. 02Data assessmentCoverage, quality, labelling and whether the target variable genuinely exists
  3. 03BaselineMeasure what the current rule or heuristic achieves on held out data
  4. 04Feasibility modelA fast, honest attempt with an explicit go or stop recommendation in writing
  5. 05Feature engineeringThe calendars, lags, ratios and groupings that carry most of the signal
  6. 06ValidationTested on time periods the model never saw, never on random splits for time series data
  7. 07DeploymentScoring pipeline, integration into the system where the decision is made, and logging of every prediction
  8. 08Monitoring and retrainingDrift alerts, scheduled retraining and a documented rollback to the previous model
DiscoverDesignBuildTestHandover
Two decisions on your side that keep the project moving

The part clients consistently underestimate is deployment. A model in a notebook is a research finding. A model that scores every new lead within a minute, logs what it predicted, survives your CRM changing a field name, and can be retrained next quarter is an engineering system. That gap is where most internal data science efforts stall, and it is a large part of what we are actually hired for.

The data questions we answer before training anything

Most machine learning projects are data projects with a modelling step at the end. We start with a hard look at what you hold: how many years, how consistently recorded, whether the definition of a field changed when you switched systems in 2022, and whether the outcome you want to predict is actually recorded anywhere. Lead scoring needs losses labelled as losses, not just quietly closed. Churn prediction needs a definition of churn that everyone agrees on before anything is trained.

Australian data has particular quirks worth handling explicitly

Australian data has particular quirks worth handling explicitly. Public holidays differ by state and shift around, school terms move demand in ways a calendar alone will not explain, the end of financial year distorts almost every B2B series in June, and weather and drought cycles matter enormously in agriculture, construction and transport. Leakage is the other classic trap: a feature that would not be available at prediction time, such as a field only populated after the sale, produces a model that looks brilliant in testing and useless in production.

  • Three or more years of history for anything with seasonality
  • Outcomes recorded honestly, including the negatives
  • Consistent field definitions across system migrations
  • State based holiday and term calendars joined in as features
  • No features that only exist after the event you are predicting
  • A written definition of the target that the business agrees on

Drift, monitoring and keeping a person in the loop

Models decay. Your product mix changes, a competitor exits the market, freight costs jump, and a model trained on last year's world starts quietly getting worse. Nobody notices for months because it still returns a number. So we log every prediction with its inputs, compare predictions against actual outcomes as they arrive, and alert when the error moves outside an agreed band. Retraining runs on a schedule with the new model compared against the incumbent before it takes over.

A score that ranks leads for a sales team is fine to automate

Where a prediction affects a person, keep the person in the loop. A score that ranks leads for a sales team is fine to automate. A score that influences credit, employment, eligibility or care needs an explanation the affected individual could be given, a route to human review, and a check for the disparities a model will happily learn from historical decisions. Using personal information to train a model is also a use under the Privacy Act 1988 that has to be consistent with the purpose it was collected for, and de-identification is worth doing where the model does not need identity to work.

When machine learning is the wrong fit

If you have two hundred records, no model will help and the honest answer is a spreadsheet and a conversation. If the question is asked once, hire an analyst instead of building a pipeline. If a simple rule captures most of the value, use the rule, because it is transparent, free to run and nobody has to maintain it. We have talked clients out of models when a well designed Power BI report answered the question outright, and out of forecasting engines when the real problem was that nobody looked at the existing forecast.

The other honest limit is organisational

The other honest limit is organisational. A model only pays off if somebody acts on its output, which usually means changing a process people have run for years. If purchasing is not going to change what they order because of the forecast, the accuracy is irrelevant. We ask that question early and in front of the people who would have to change. Where the need is really visibility rather than prediction, analytics implementation delivers value faster, and for logistics operators in particular clean measurement usually has to come first anyway.

How we scope it

Four ways to scope your Machine Learning project

We do not publish package prices, because the same brief can be a short build or a long one. These are the shapes the work usually takes. Tell us which one sounds like you and you will get a fixed written quote that spells out exactly what it covers.

Proof of value

One use case, evaluated honestly before it goes near a customer

Fixed written quote, agreed before work starts

  • Data assessment with quality issues quantified
  • Documented baseline the model must beat
  • Feasibility report with a written go or stop recommendation
Request a quote
Most common

Production build

In production, with a human approval step and an evaluation set

Fixed written quote, agreed before work starts

  • Everything in Proof of value
  • Trained model with validation on unseen time periods
  • Feature definitions documented for reproducibility
  • Scoring pipeline integrated where the decision is made
Request a quote

Embedded platform

Built into the product rather than bolted onto it

Fixed written quote, agreed before work starts

  • Everything in Production build
  • Prediction logging with outcome comparison
  • Drift monitoring, alerting and a retraining schedule
  • Handover documentation and rollback procedure
Request a quote

Model care

Monitoring, evaluation and retraining as the inputs drift

Rolling monthly, quoted in writing

  • Evaluation set rerun as the model and the inputs change
  • Cost and quality reported monthly, not assumed
  • Prompt, tool and guardrail changes as the work shifts
  • Rolling, cancel with 30 days notice
Request a quote

These are shapes, not menus. Most quotes end up somewhere between two of them, and we will say so when the honest answer is the smallest one. Describe the problem and we will tell you which it is.

Questions buyers usually ask

Frequently asked questions

Scope and timeline

How long does a machine learning project take?

Usually 12 to 24 weeks end-to-end, though the feasibility phase is much shorter and gives you a decision point after a limited spend. Data preparation typically consumes more of the schedule than modelling. If your records need cleaning or unifying across systems first, that work is scoped separately so you can see what it is costing you.

What happens when the model starts getting worse?

You find out from monitoring rather than from a complaint. Predictions are compared against actual outcomes continuously and an alert fires when error exceeds the agreed band. Retraining is scheduled, and each candidate model is compared against the one in production before replacing it. If the underlying business has changed shape, the fix may be new features rather than more data, and we look at that together.

Detail and edge cases

What drives the cost of a machine learning build?

Data condition first, then the number of models, the integration work to put predictions where decisions are made, and the level of monitoring required. Running costs are usually modest compared with language model work, since a trained model is cheap to run. We quote in writing after the data assessment, and we quote the feasibility phase separately so you can stop there.

Do we own the model and the training pipeline?

Yes. The code, the pipeline, the feature definitions and the model artefacts are yours, held in your repository and running in your cloud accounts. Your data is never pooled with another client's or used to train anything outside your project. If you take it in-house, the documentation is written for a competent data engineer who has never met us.

How much data do we actually need?

It depends on the question, but useful rules of thumb exist. Seasonal forecasting wants three or more years so each season appears several times. Classification wants at least several hundred examples of the rarer outcome, not just of the common one. Thin data is not automatically fatal, though it does mean simpler models and wider uncertainty, and we will say so plainly.

Can machine learning help with demand and inventory planning?

It is one of the strongest use cases, particularly for businesses carrying stock across multiple states where freight and lead times differ. Forecasts feed reorder points rather than replacing the planner, and a person still reviews unusual recommendations. Value comes from the tail: better calls on slow moving lines and seasonal peaks. It pairs naturally with an inventory system that already holds clean movement data.

Find out whether your data can answer the question

Tell us the decision you want to improve and roughly what history you hold. We reply within one business day, and the feasibility phase is quoted separately so you can stop if the answer is no.