The request arrives as one line: make the order system available to the AI assistant. The team already runs a REST API, so the first decision is MCP vs API.

An MCP server earns its place when an agent must choose among your system's operations at run time; for a known call from your own code, the existing API is simpler. Either way, most of the work is deciding whose permissions the call runs with.

MCP vs API: what the protocol adds over a REST endpoint

Anthropic open-sourced the Model Context Protocol on 25 November 2024 and donated it to the Agentic AI Foundation, a fund under the Linux Foundation, in December 2025. The specification defines three roles: a host (the AI application), the client inside it that holds a connection, and a server that offers capabilities. Messages are JSON-RPC 2.0.

A server offers three kinds of thing. Tools are operations the model may call, each with a name, a description in plain language and a JSON Schema for its input. Resources are data the application can read into the model's context, and prompts are templates a user can pick.

The difference from a REST API is who reads the contract. An OpenAPI document is written for a developer, who reads it once and writes the calls into code. An MCP server describes itself to a model at run time, and the model decides which tool to call and with what arguments. The tool description is part of the prompt, which is why it is also an attack surface.

The July 2026 revision of the specification narrowed the gap on the operational side. It removed protocol-level sessions and the initialize handshake: MCP is now stateless, every request carries its protocol version and the client's capabilities, and a server that needs state across calls returns a handle the client passes back as an ordinary tool argument. An MCP server over Streamable HTTP now scales behind a load balancer like any other stateless service.

REST APIMCP server
Written forA developer who writes the callsA model that picks the calls at run time
ContractOpenAPI, read at build timeTool names, descriptions and JSON Schemas, listed at run time
Who decides what is calledYour codeThe model, within the tools it is given
StateUsually statelessStateless since the 2026-07-28 revision; handles passed as arguments
AuthorizationWhatever you builtOAuth 2.1 profile in the specification, optional but expected over HTTP
TestingDeterministic: same call, same resultThe service is deterministic; the choice of call is not

MCP vs function calling

Function calling (tool use) is a feature of a model provider's API. Your application sends tool definitions with each request, the model answers with the call it wants, and your code runs it. In Java, Spring AI and LangChain4j both turn an annotated method into such a tool, inside your own process.

MCP standardizes where those tool definitions come from. The same server works with any MCP client: a chat assistant, a coding agent in an IDE, or a Spring AI application of your own. If one application talks to one model and calls a handful of your methods, function calling inside that application is enough. MCP starts to pay when several clients need the same capabilities and you want one place to control them.

When to put an MCP server in front of a Java application, and when not to

Matteo Rossi's InfoQ article on MCP in Java (April 2026) makes the architectural case: "MCP servers act as anti-corruption layers between LLMs and core systems, exposing controlled capabilities rather than raw APIs." He is as clear about the price. "For small teams or short-lived experiments, these costs likely outweigh the benefits", the costs being extra components, network hops and more points of failure.

The table applies that trade-off case by case.

SituationUseWhy
Your own backend calls a known endpoint in a fixed sequenceThe REST API you haveNo model in the loop, so nothing to discover; deterministic and easy to test
One AI feature in one application, a few of your methodsFunction calling in-processNo extra service to deploy, secure and monitor
Several AI clients need the same capabilities (an assistant, IDE agents, a partner's agent)An MCP serverOne contract and one place for authorization and audit
An agent has to choose among many operations depending on the conversationAn MCP serverDiscovery and schemas are what the protocol is for
The operation moves money or changes records that cannot be undoneAn API behind a workflow with human approval, or a tool that only draftsThe model proposes; a person or a fixed rule commits
A short experimentFunction callingThe overhead of a separate server is not repaid

Whichever row you land on, design the tools as business capabilities. find_open_orders(customerId) is a better tool than a generic GET /orders with a free-form filter: it is easier for the model to choose correctly, and easier for you to authorize and log.

Building an MCP server in Java with Spring AI

The official MCP Java SDK reached version 2.0.0 in June 2026, and Spring AI's MCP support extends it with Spring Boot starters and annotations. Version 2.0.0 of the SDK tracks the 2025-11-25 revision of the specification, so check which revision your SDK speaks before relying on the stateless behavior of July 2026.

With the Spring AI server starter, a tool is an annotated method on a Spring bean. Keep it thin: the @McpTool method calls the same service your REST controllers call, so validation and business rules stay in one place.

import org.springframework.ai.mcp.annotation.McpTool;
import org.springframework.ai.mcp.annotation.McpToolParam;

@Component
public class OrderTools {

    private final OrderService orders;   // the service your REST controllers already call

    public OrderTools(OrderService orders) {
        this.orders = orders;
    }

    @McpTool(name = "find_open_orders",
             description = "Open orders for one customer, newest first, at most 20",
             annotations = @McpTool.McpAnnotations(readOnlyHint = true, destructiveHint = false))
    public List<OrderSummary> findOpenOrders(
            @McpToolParam(description = "Customer number, for example C-10442", required = true)
            String customerId) {
        return orders.findOpen(customerId, 20);
    }
}
# spring-ai-starter-mcp-server-webmvc on the classpath
spring.ai.mcp.server.protocol=STATELESS
spring.ai.mcp.server.annotation-scanner.enabled=true

Without the annotations line, Spring AI 2.0.1 advertises every tool as destructive and not read-only, so a client that asks before risky calls would treat a lookup like a delete. The hints only inform the client: the specification says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers", and the permission check still belongs in the service.

The starter offers STREAMABLE and STATELESS for HTTP, and STDIO for a server that runs as a local process. For a server that several applications reach over the network, stateless HTTP is the one that scales and fails like the rest of your services.

The Spring documentation warns that the HTTP transports "expose an unauthenticated JSON-RPC endpoint by default", and "any client that can reach the endpoint can list and invoke every registered tool". Security is yours to add, with Spring Security or the Spring AI MCP Security module, which its documentation still marks as work in progress.

MCP authorization: whose permissions the agent uses

Authorization is optional in the MCP specification, but an HTTP server that implements it should follow the specification's OAuth 2.1 profile. In that profile the MCP server is an OAuth resource server. It publishes Protected Resource Metadata (RFC 9728) so clients can find the authorization server, and clients must name the server they want a token for with a resource indicator (RFC 8707).

The specification is strictest about tokens. The server must validate that a token was issued for it as the audience, and the specification says: "MCP servers MUST NOT accept or transit any other tokens." Forwarding the client's token to your downstream API is called token passthrough, and it is forbidden because the downstream service can no longer tell who is calling, and a stolen token turns your MCP server into a proxy for someone else.

In a Spring application that comes down to a few decisions:

  • Configure the MCP server as a Spring Security resource server and check the aud claim as well as the signature.
  • Give tools scopes that match the risk: orders:read for lookups, a separate orders:write for anything that changes data. The specification's security guidance lists publishing every scope up front and wildcard scopes such as * as common mistakes, and expects a 403 with insufficient_scope when a tool needs more.
  • Call your own services with the MCP server's credentials plus the user's identity, through a token exchange or an internal header your gateway trusts, and let the order service apply the same permissions it applies to the web application.
  • Log the user, the client, the tool and the arguments of every call, so an action taken by an agent can be traced to the person it acted for.
The request path from an AI application to your data. The AI application, acting as MCP client, asks for a token issued for this MCP server only. The MCP server, a thin layer, rejects tokens issued for anything else, grants one scope per tool, never forwards the client's token and logs every call. It calls the existing order service, which checks the user's own permissions and validates input as it does for any client. The database is never reached by the agent directly.CALLERNEW LAYERWHAT YOU ALREADY RUNAI applicationMCP clientMCP servera thin layerOrder serviceyour existing JavaDatabaseunchangedAsks for a tokenfor this server only(resource indicator)Rejects tokensissued for othersOne scope per toolNever forwardsthe client's tokenLogs every callChecks the user'sown permissionsValidates input asfor any clientNever reached bythe agentdirectly
Where the checks sit when an MCP server fronts an existing Java service. The server adds token and scope checks; the service keeps the permission rules it already has.

A server that runs over STDIO on a developer's machine is a different case: the specification says it should take credentials from the environment instead. That is fine for a personal tool and not for a server several people share.

MCP security: tool poisoning, prompt injection and the incidents so far

Three cases reported by security vendors in 2025 show the main risks. In the first two, the model treats text that came from outside as an instruction, and a tool with broad access carries it out. The third is an ordinary injection bug in a client.

  • Tool poisoning. In April 2025 Invariant Labs showed that a malicious server can hide instructions in a tool description, telling the model to read files such as SSH keys and send them out, while the user sees an ordinary-looking tool.
  • A poisoned GitHub issue. In May 2025 Invariant Labs reported that a malicious issue in a public repository could make an agent using the official GitHub MCP server leak data from the user's private repositories. Invariant stressed that this was not a flaw in the server code: the agent's token reached more repositories than the task needed.
  • A client-side remote code execution. In July 2025 JFrog disclosed CVE-2025-6514 (CVSS 9.6) in mcp-remote, a popular proxy for MCP clients: connecting to an untrusted server could run operating system commands on the user's machine.

The specification's security best practices now cover these classes of problem, from confused deputy attacks on proxy servers to the new state handle hijacking: since MCP became stateless, a server must bind every handle it issues to the authenticated user and never treat possession of a handle as proof of identity.

MCP security checklist for a Java team

  • Put authentication in front of every HTTP transport before the server leaves localhost.
  • Accept only tokens issued for this server, and never forward them downstream.
  • Scope per tool, with write and delete operations in separate scopes from reads.
  • Let the existing service apply the user's permissions; the MCP layer does not replace them.
  • Treat every tool result that contains outside text (emails, issues, documents) as untrusted input to the model.
  • Require a human approval step for operations that move money or cannot be undone.
  • Bind state handles to the user server-side, generate them with a secure random generator, and let them expire.
  • Install third-party MCP servers only from sources you have reviewed, pinned to a version.
  • Log user, client, tool and arguments for every call, and alert on unusual volumes.
  • Cap the number of tool calls per request, so a looping agent stops on its own.

MCP gateways: when one is worth adding

An MCP gateway is a proxy in front of several MCP servers that gives them one entry point for authentication, tool allowlists, rate limits and logging. It earns its place when a company runs several MCP servers for several AI clients and wants one policy across all of them.

With one server, Spring Security inside that server does the same job with one component fewer. A gateway also becomes a proxy in the specification's sense, so the confused deputy rules apply to it: per-client consent before it forwards anyone to a third-party authorization server.

What the LangChain4j experiment shows about agents calling your tools

Kevin Dubois and Mario Fusco had a coding assistant build an agent with LangChain4j and ran it in two shapes. In the supervisor pattern one agent decides at run time which sub-agent to call next; in the workflow pattern the sequence is fixed in code. Both fixed the bugs they were given. The workflow finished in two minutes, against more than six for the supervisor, and the authors put the difference down to the supervisor's own coordination overhead.

Their first run also hit a limit worth knowing. With gpt-4o, the model "got stuck in a tool-calling loop, which LangChain4j interrupted" after the default maximum of 100 sequential tool invocations.

Where the steps are known, a fixed workflow calling your tools over MCP is faster and easier to test than an agent choosing them. The client also needs a hard cap on tool calls, because your services will see every call of a looping agent, with the rate limits and audit trail that implies.

How we approach an MCP integration

We start an AI integration from the use case and choose the protocol last. In a four-week pilot we agree which operations the model may call, under whose permissions, with an evaluation set that shows whether it chooses them correctly, and we decide between function calling, an existing API and an MCP server on that evidence. The tools call the Java services you already run, so the rules your web application enforces apply to the agent too.

If you have been asked to make a Java system available to an AI assistant, that is what our AI Integration pilot starts from: a thirty-minute call with the architect who would lead the work, and a written ballpark within a week.