FAQ

Frequently asked questions

Questions I hear from prospects evaluating local or hybrid AI -- and the honest answers.

How It Works

What does "local AI" actually mean?

Local or client-hosted inference means the AI model runs in an environment controlled or approved by you rather than through a provider-hosted model endpoint like OpenAI or Google. This reduces third-party data exposure, but local inference describes where the model executes -- it does not guarantee that no data ever leaves your environment. External integrations, telemetry, storage, and support access are mapped separately for each project.

So nothing touches the cloud?

It depends on the design. Some systems are fully local -- inference, memory, and tools all run on client-controlled hardware. Others use a hybrid approach: sensitive processing stays local while specific integrations (email delivery, market data feeds, telephony) connect to managed services. Some projects also use managed model APIs when the client approves that data flow and integration, latency, or operating requirements make it the better fit. I match the architecture to what each workflow actually requires.

What's the difference between local, hybrid, and cloud-connected?

Client-hosted
The AI model runs on hardware you control. Your sensitive data stays within your network for processing, though other parts of the system (email, databases, APIs) may still connect to external services as the workflow requires.
Hybrid
Model inference is local, but the system integrates with external services that the workflow depends on. Different stages of a pipeline may run in different approved environments.
Cloud-connected
The application uses managed cloud services, which may include provider-hosted models. This is appropriate when integration, latency, or operating requirements outweigh the need for on-premise processing.

Hardware & Costs

Do I need to buy my own GPU?

Not necessarily. Hardware requirements depend on the model chosen, quantization level, context length, concurrency needs, and latency target. Some low-volume workloads can run on CPUs or a single consumer-grade GPU; production workloads may require more VRAM, redundancy, or managed infrastructure. During discovery, I benchmark your actual workflow against representative data before recommending hardware -- many SMBs find their requirements are lower than expected, but the bar varies significantly by use case.

How much does local AI cost to run?

It varies widely by scale and workload. A single GPU workstation costs roughly $2,000-$5,000 upfront (consumer) while enterprise-grade systems can range from $10,000 to well over $30,000 depending on whether you need redundancy, ECC memory, or professional support contracts. Electricity cost depends on measured wall power, utilization, cooling overhead, and local rates. Moving inference locally may avoid hosted-model token charges for the workload moved on-premise, though managed integrations and other API costs often remain. I compare total cost of ownership -- hardware, electricity, maintenance, support, and external services -- against cloud pricing before recommending a change.

Can we start without buying hardware?

Yes. I build and demo prototypes on my own infrastructure first, so you see the system work with representative data before committing to hardware or full deployment. The production target -- whether it's your hardware, a private cloud instance, or a hybrid setup -- is agreed during discovery.

What happens if our model needs to grow?

I isolate model-specific integration where practical so models can be evaluated or replaced without rebuilding the entire system. A model change may still require different hardware capacity, runtime adjustments, prompt updates, and revalidation against your data. Larger models are also not automatically better for every task -- I test each candidate against your actual workflow before recommending a swap.

Data & Compliance

Does running AI locally make us HIPAA-compliant?

No. Local inference reduces third-party data exposure, but it does not by itself establish compliance. HIPAA involves risk analysis and administrative, physical, and technical safeguards that go beyond where a model runs. Depending on the parties' roles and whether I or another vendor will create, receive, maintain, or transmit PHI, appropriate agreements such as a BAA may be required before access is provided. HHS also expressly allows compliant cloud use when appropriate safeguards are in place -- local deployment is one option among several. I implement scoped technical controls; the client remains responsible for its compliance program with advice from counsel and security professionals.

What about SOC 2 or GDPR?

Local deployment is an architectural option, not a compliance certification. SOC 2 requires an independent CPA examination of controls at a service organization -- it's not something a single consultant can confer. GDPR is law; the data controller retains responsibility for demonstrating compliance with processing obligations. I can design systems that support compliance efforts -- audit trails, access logs, technical retention controls implementing your approved policy -- but neither a security team nor an auditor becomes the final legal decision-maker for your organization's obligations.

Is my data safe during the prototyping phase?

Prototypes typically use synthetic, de-identified, or otherwise approved representative data in an agreed environment. If identifiable or otherwise sensitive production data is necessary for a meaningful demo, it is used only after the required contractual, security, and data-handling terms are documented and approved by you before work begins.

Do you store my data after the project ends?

Project data is returned or deleted according to the engagement's agreed retention schedule, subject to applicable legal, contractual, security, and backup-retention requirements. Deliverables (code, documentation, deployment configs) are handed over as defined in the project scope. The proposal identifies the deliverables, handoff process, and any records that must be retained beyond active engagement.

Engagement & Process

How long does a typical project take?

It depends on scope:

Workflow Fit Assessment
~2 weeks. I review your infrastructure and processes, then deliver a concrete plan with cost estimates.
Private AI Pilot
2-3 weeks from kickoff to working prototype you can evaluate with representative data.
AI Pipeline Development
3-6 weeks depending on integration complexity and testing requirements.
Production Deployment
4-12 weeks depending on environment preparation, security review, and acceptance criteria.

Complex integrations, client environment readiness, or extensive security review may extend these timelines. The proposal states the applicable timeline for your specific project.

What if we're not technical -- do we need an in-house AI team?

No. I design systems for handoff to designated staff and define the required documentation, training, runbook, and support responsibilities in each proposal. For ongoing support, my monthly retainer covers a defined allocation of troubleshooting, model updates, maintenance tasks, and consulting within agreed hours and scope. Whether you need to hire full-time depends on your workflow volume and internal capacity.

What if the AI gives wrong answers?

All language models -- local or cloud -- can produce incorrect outputs. I design systems with guardrails: retrieval-augmented generation (RAG) improves grounding by referencing your own documents, output validation catches errors where outputs can be checked against schemas or rules, and human review steps are defined for high-stakes decisions. RAG alone cannot guarantee correctness -- retrieval may return incomplete or outdated content. The goal is to design systems that reduce and manage risk, not eliminate it.

Can you work with our existing IT vendor?

I can coordinate with your IT provider or internal technical team. Integration with existing authentication, monitoring, backup, and security infrastructure is assessed during discovery and included where supported and within project scope. I'm adding AI capability to your stack, not replacing your infrastructure team.

How do you price your work?

Proposals use fixed-price or milestone-based structures depending on scope. Ongoing advisory work starts at $3,000 per month with a defined allocation and support boundaries. During the initial strategy call -- which is free and carries no obligation -- we discuss whether your project fits a fixed-price or retainer model. The proposal states the price, assumptions, exclusions, milestones, and change process before work begins.

Did not find what you were looking for?

Book a Free Strategy Call