It’s been a while. I’m about to refactor the DA app server, and one of the things I’d like to explore is adding Amazon Bedrock.
My first thought is to build a Bedrock service directly into the DA app server. What I’m not sure about yet is the best way to structure it so that we can use it at the app-server level while also making the appropriate functionality available to clients. I’d like to get your opinion on how you would approach that.
The other option I’ve been kicking around is putting the AI functionality in its own application/service and having the DA app server interact with it. That might give us better visibility, separation, and control, but it also makes the architecture more complex.
I’m mostly spitballing here, so I’m definitely open to a completely different approach. Let me know what you think and if there are any pitfalls you think I should be aware of.
First, a small plain .NET interface that calls Bedrock. No RO attributes, no
service class.
Second, a class on top of it that takes your data and builds the prompt. Your
server-side code calls that class directly. The RO service calls the same
class. One implementation, used by both.
For the client side, I would not give clients a generic method. Instead of
Ask(prompt), something like SummarizeOrder(orderId). The client sends an id,
and the server builds the prompt from its own data. Then the client cannot
change the prompt or pick the model, and you stay in control of the cost.
A separate process is just another implementation of the same interface,
chosen by configuration. I would keep it inside the app server first.
Attached is the test project I used. It has both versions in it so you can
compare them
It runs without AWS using a stub, so you can set your own delay and try your
own. scenario BedrockAppServer.zip (32.4 KB)
The app server holds a worker thread until method finishes executing.
While a database call takes milliseconds, a call to Bedrock takes seconds.
I sent a lot of AI requests at once and the server ran out of worker threads.
That blocked everything else too - even a simple method that just returns a
number took over 10 seconds to respond.
Clients that don’t use AI were stuck waiting as well.