fn-serverless

mirror of https://github.com/fnproject/fn.git synced 2022-10-28 21:29:17 +03:00

Author	SHA1	Message	Date
Reed Allman	c9e995292c	if a slot is available, don't launch more (#701 ) since we were sending a signal before checking if a slot was available, even in the case of serial calls locally I was seeing 2 containers launch. if we only send a signal after first checking if a slot is available, this goes away. 1 usec should not be too offensive of an additional wait, all things considered here.	2018-01-18 13:19:25 -08:00
Tolga Ceylan	5a7778a656	fn: cancellations in WaitAsyncResource (#694 ) * fn: cancellations in WaitAsyncResource Added go context with cancel to wait async resource. Although today, the only case for cancellation is shutdown, this cleans up agent shutdown a little bit. * fn: locked broadcast to avoid missed wake-ups * fn: removed ctx arg to WaitAsyncResource and startDequeuer This is confusing and unnecessary.	2018-01-17 16:08:54 -08:00
Nigel Deakin	8bf26efa29	Add new Prom metrics fn_timeout and fn_errors (#679 ) * Add new Prom metric fn_timedout * Add new Prometheus metric fn_errors * Tidy up variable name * Add new Prometheus metric fn_errors * gofmt	2018-01-15 14:49:33 +00:00
Reed Allman	0bde666395	clean up agent.Submit (#681 ) this was getting bloated with various contexts and spans and stats administrivia that obfuscated what was going on a lot. this makes some helper methods to shove most of that stuff into, and simplifies the context handling around getting a slot by moving it inside of slot acquisition code. also removed most uses of `call.Model()` -- I'll kill this thing some day, but if a reason is needed, then the overhead of dynamic dispatch is unnecessary, we're inside of the implementee for the agent, we don't want to use the interface methods inside of that.	2018-01-12 13:56:17 -08:00
Tolga Ceylan	39b2cb2d9b	Cpu resources (#642 ) * fn: cpu quota implementation	2018-01-12 11:38:28 -08:00
Tolga Ceylan	db159e595f	fn: new container lauch adjustments (#677 ) ) revert executor wait queue size comparison. This is too aggresive and with stall check below, now unnecessary. ) new container logic now checks if stats are constant, if this is the case, then we assume the system is stalled (eg running functions that take long time), this means we need to make progress and spin up a new container.	2018-01-11 14:09:21 -08:00
Nigel Deakin	ac2bfd3462	Change basic stats to use opentracing rather than Prometheus API (#671 ) * Change basic stats to use opentracing rather than Prometheus API directly * Just ran gofmt * Extract opentracing access for metrics to common/metrics.go * Replace quotes strings with constants where possible	2018-01-11 17:34:51 +00:00
Reed Allman	20089c4e83	make headers quasi-consistent (#660 ) possible breakages: * `FN_HEADER` on cold are no longer `s/-/_/` -- this is so that cold functions can rebuild the headers as they were when they came in on the request (fdks, specifically), there's no guarantee that a reversal `s/_/-/` is the original header on the request. * app and route config no longer `s/-/_/` -- it seemed really weird to rewrite the users config vars on these. should just pass them exactly as is to env. * headers no longer contain the environment vars (previously, base config; app config, route config, `FN_PATH`, etc.), these are still available in the environment. this gets rid of a lot of the code around headers, specifically the stuff that shoved everything into headers when constructing a call to begin with. now we just store the headers separately and add a few things, like FN_CALL_ID to them, and build a separate 'config' now to store on the call. I thought 'config' was more aptly named, 'env' was confusing, though now 'config' is exactly what 'base_vars' was, which is only the things being put into the env. we weren't storing this field in the db, this doesn't break unless there are messages in a queue from another version, anyway, don't think we're there and don't expect any breakage for anybody with field name changes. this makes the configuration stuff pretty straight forward, there's just two separate buckets of things, and cold just needs to mash them together into the env, and otherwise hot containers just need to put 'config' in the env, and then hot format can shove 'headers' in however they'd like. this seems better than my last idea about making this easier but worse (RIP). this means: * headers no longer contain all vars, the set of base vars can only be found in the environment. * headers is only the headers from request + call_id, deadline, method, url * for cold, we simply add the headers to the environment, prepending `FN_HEADER_` to them, BUT NOT upper casing or `s/-/_/` * fixes issue where async hot functions would end up with `Fn_header_` prefixed headers * removes idea of 'base' vars and 'env'. this was a strange concept. now we just have 'config' which was base vars, and headers, which was base_env+headers; i.e. they are disjoint now. * casing for all headers will lean to be `My-Header` style, which should help with consistency. notable exceptions for cold only are FN_CALL_ID, FN_METHOD, and FN_REQUEST_URL -- this is simply to avoid breakage, in either hot format they appear as `Fn_call_id` still. * removes FN_PARAM stuff * updated doc with behavior weird things left: `Fn_call_id` e.g. isn't a correctly formatted http header, it should likely be `Fn-Call-Id` but I wanted to live to fight another day on this one, it would add some breakage. examples to be posted of each format below closes #329	2018-01-09 10:08:30 -08:00
Tolga Ceylan	18716911b9	fn: agent slot and execution wait correction (#658 ) Since by policy we require timeout/2 remaining time before we can execute the request, we should also bound the slot wait time by timeout/2 to avoid waiting for full timeout in slot wait phase.	2018-01-08 12:33:37 -08:00
Tolga Ceylan	14789aba41	Slot mgr fixes (#613 ) ) during shutdown, errors should be 503 ) new inactivity time out for hot queue, we previously kept hot queues in memory forever. ) each hot queue now has a hot launcher to monitor and launch hot containers ) consumers now create a consumer channel with startDequeuer() that can be cancelled via context ) consumers now ping (signal) hot launcher every 200 msecs until they get a slot ) tests for slot queue & mgr	2018-01-04 11:34:43 -08:00
Tolga Ceylan	feeeca3321	fn: agent shutdown improvements (#622 )	2017-12-22 12:52:31 -08:00
Tolga Ceylan	25a72146f5	slot tracking improvements (#562 ) * fn: remove 100 msec sleep for hot containers ) moved slot management to its own file ) slots are now implemented with LIFO semantics, this is important since we do not want to round robin hot containers. Idle hot containers should timeout properly. ) each slot queue now stores a few basic stats such as avg time a call spent in a given state and number of running/launching containers, number of waiting calls in those states. ) first metrics in these basic stats are discarded to avoid initial docker pull/start spikes. ) agent now records/updates slot queue state and how much time a call stayed in that state. ) waitHotSlot() replaces the previous wait 100 msec logic where it sends a msg to hot slot go routine launchHot() and waits for a slot *) launchHot() is now a go routine for tracking containers in hot slots, it determines if a new containers is needed based on slot queue stats.	2017-12-15 15:50:07 -08:00
Nigel Deakin	f1fc040948	Fix spans for prometheus (#606 )	2017-12-15 10:31:57 -08:00
Tolga Ceylan	eccce881a6	fn: exclude timeouts from failed error count (#590 ) * fn: exclude timeouts from failed error count	2017-12-14 13:10:07 -08:00
Reed Allman	bb92547b95	Hybrid plumby (#585 ) * fix configuration of agent and server to be future proof and plumb in the hybrid client agent * fixes up the tests, turns off /r/ on api nodes * fix up defaults for runner nodes * shove the runner async push code down into agent land to use client * plumb up async-age * return full call from async dequeue endpoint, since we're storing a whole call in the MQ we don't need to worry about caching of app/route [for now] * fast safe shutdown of dequeue looper in runner / tidying of agent * nice errors for path not found against /r/, /v1/ or other path not found * removed some stale TODO in agent * mq backends are only loud mouths in debug mode now * update tests * Add caching to hybrid client * Fix HTTP error handling in hybrid client. The type switch was on the value rather than a pointer. * Gofmt. * Better caching with a nice caching wrapper * Remove datastore cache which is now unused * Don't need to manually wrap interface methods * Go fmt	2017-12-12 15:54:55 -08:00
Reed Allman	2ebc9c7480	hybrid mergy (#581 ) * so it begins * add clarification to /dequeue, change response to list to future proof * Specify that runner endpoints are also under /v1 * Add a flag to choose operation mode (node type). This is specified using the `FN_NODE_TYPE` environment variable. The default is the existing behaviour, where the server supports all operations (full API plus asynchronous and synchronous runners). The additional modes are: * API - the full API is available, but no functions are executed by the node. Async calls are placed into a message queue, and synchronous calls are not supported (invoking them results in an API error). * Runner - only the invocation/route API is present. Asynchronous and synchronous invocation requests are supported, but asynchronous requests are placed onto the message queue, so might be handled by another runner. * Add agent type and checks on Submit * Sketch of a factored out data access abstraction for api/runner agents * Fix tests, adding node/agent types to constructors * Add tests for full, API, and runner server modes. * Added atomic UpdateCall to datastore * adds in server side endpoints * Made ServerNodeType public because tests use it * Made ServerNodeType public because tests use it * fix test build * add hybrid runner client pretty simple go api client that covers surface area needed for hybrid, returning structs from models that the agent can use directly. not exactly sure where to put this, so put it in `/clients/hybrid` but maybe we should make `/api/runner/client` or something and shove it in there. want to get integration tests set up and use the real endpoints next and then wrap this up in the DataAccessLayer stuff. * gracefully handles errors from fn * handles backoff & retry on 500s * will add to existing spans for debuggo action * minor fixes * meh	2017-12-11 10:43:19 -08:00
Tolga Ceylan	9481f811b7	fn: fail count should include timeouts (#577 ) * fn: fail count should include timeouts	2017-12-06 16:11:59 -08:00
Nigel Deakin	96f27070be	More metrics (#561 ) * Add new spans to agent.submit * Add new spans to agent.submit * Add new spans to agent.submit * Add new spans to agent.submit	2017-12-05 10:26:28 -08:00
Travis Reeder	0798f9fac8	Middleware upgrade (#554 ) * Adds root level middleware * Added todo * Better way for extensions to be added. * Bad conflict merge?	2017-12-05 08:22:03 -08:00
Tolga Ceylan	25f6706642	Container memory tracking related changes (#541 ) * squash# This is a combination of 10 commits2 fn: get available memory related changes ) getAvailableMemory() improvements ) early fail if requested memory too large to meet ) tracking async and sync pools individually. Sync pool is reserved for sync jobs only, while async pool can be used by all jobs. ) head room estimation for available memory in Linux.	2017-12-01 11:21:16 -08:00
Reed Allman	892c843d87	add error to call model (#539 ) * add error to call model closes #331 previously, for async this error was being masked completely even if it was something useful like the image not existing. for sync, the error was returned in the http request but now it's also being stored. this error itself can cover a lot of landscape, it could be an error in getting a slot, pulling an image, running a container, among other things. anyway, no longer being masked. we can likely improve it in certain cases we run into in the future, but it's open ended at the moment and not being masked like some errors in sync http request returns (503 non-models.APIError) for now. * tucks in callTrigger stuff to keep api clean * adds swagger * adds migration * adds tests for datastore and agent to ensure behavior * pull images before tests are ran * gofmt migrations file	2017-11-28 11:21:39 -06:00
Nigel Deakin	954f69e74a	Add appname to basic metrics (#547 ) * Add app labels to queued/running/completed/failed metrics * Add app labels to queued/running/completed/failed metrics * Add app labels to queued/running/completed/failed metrics	2017-11-28 10:17:24 -06:00
Reed Allman	c9198b8525	add per call stats field as histogram (#528 ) * add per call stats field as histogram this will add a histogram of up to 240 data points of call data, produced every second, stored at the end of a call invocation in the db. the same metrics are also still shipped to prometheus (prometheus has the not-potentially-reduced version). for the API reference, see the updates to the swagger spec, this is just added onto the get call endpoint. this does not add any extra db calls and the field for stats in call is a json blob, which is easily modified to add / omit future fields. this is just tacked on to the call we're making to InsertCall, and expect this to add very little overhead; we are bounding the set to be relatively small, planning to clean out the db of calls periodically, functions will generally be short, and the same code used at a previous firm did not cause a notable db size increase with production workload that is worse, wrt histogram size (I checked). the code changes are really small aside from changing to strfmt.DateTime, adding a migration and implementing sql.Valuer; needed to slightly modify the swap function so that we can safely read `call.Stats` field to upload at end. with the full histogram in hand, we can compute max/min/average/median/growth rate/bernoulli distributions/whatever very easily in a UI or tooling. in particular, this data is easily chartable [for a UI], which is beneficial. * adds swagger spec of api update to calls endpoint * adds migration for call.stats field * adds call.stats field to sql queries * change swapping of hot logger to exec, so we know that call.Stats is no longer being modified after `exec` [in call.End] * throws out docker stats between function invocations in hot functions (no call to store them on, we could change this later for debug; they're in prom) * tested in tests and API closes #19 * add format of ints to swag	2017-11-27 08:52:53 -06:00
Tolga Ceylan	2551be446a	fn: introducing 503 responses for out of capacity case (#518 ) * fn: introducing 503 responses for out of capacity case ) Adding 503 with Retry-After header case if request failed during waiting for slots. ) TODO: return 503 without Retry-After if the request can never be met by this fn server. ) fn: runner test docker pull fixup ) fn: MaxMemory for routes is now a variable to allow testing and adjusting it according to fleet memory sizes.	2017-11-21 12:42:02 -08:00
Reed Allman	2d8c528b48	S3 loggyloo (#511 ) * add minio-go dep, update deps * add minio s3 client minio has an s3 compatible api and is an open source project and, notably, is not amazon, so it seems best to use their client (fwiw the aws-sdk-go is a giant hair ball of things we don't need, too). it was pretty easy and seems to work, so rolling with it. also, minio is a totally feasible option for fn installs in prod / for demos / for local. * adds 's3' package for s3 compatible log storage api, for use with storing logs from calls and retrieving them. * removes DELETE /v1/apps/:app/calls/:call/log endpoint * removes internal log deletion api * changes the GetLog API to use an io.Reader, which is a backwards step atm due to the json api for logs, I have another branch lined up to make a plain text log API and this will be much more efficient (also want to gzip) * hooked up minio to the test suite and fixed up the test suite * add how to run minio docs and point fn at it docs some notes: notably we aren't cleaning up these logs. there is a ticket already to make a Mr. Clean who wakes up periodically and nukes old stuff, so am punting any api design around some kind of TTL deletion of logs. there are a lot of options really for Mr. Clean, we can notably defer to him when apps are deleted, too, so that app deletion is fast and then Mr. Clean will just clean them up later (seems like a good option). have not tested against BMC object store, which has an s3 compatible API. but in theory it 'just works' (the reason for doing this). in any event, that's part of the service land to figure out. closes #481 closes #473 * add log not found error to minio land	2017-11-20 17:39:45 -08:00
Tolga Ceylan	17d4271ffb	fn: move memory/token code into resource (#512 ) ) bugfix: fix nil ptr access in docker registry RoundTrip ) move async and ram token related code into resource.go	2017-11-17 15:25:53 -08:00
Nigel Deakin	910612d0b1	Docker stats to Prometheus (#486 ) * Docker stats to Prometheus * Fix compilation error in docker_test * Refactor docker driver Run function to wait for the container to have stopped before stopping the colleciton of statistics * Fix go fmt errors * Updates to sending docker stats to Prometheus * remove new test TestWritResultImpl because we changes to support multiple waiters have been removed * Update docker.Run to use channels not contextrs to shut down stats collector	2017-11-16 11:02:33 -08:00
Travis Reeder	96cfc9f5c1	Update json (#463 ) * wip * wip * Added more fields to JSON and added blank line between objects. * Update tests. * wip * Updated to represent recent discussions. * Fixed up the json test * More docs * Changed from blank line to bracket, newline, open bracket. * Blank line added back, easier for delimiting.	2017-11-16 09:59:13 -08:00
Tolga Ceylan	a530cd9be3	Minor naming and control flow changes to satisfy golint	2017-11-02 15:36:55 -07:00
Reed Allman	ce252d0448	Merge pull request #424 from fnproject/call-listener CallListener - replaces RunnerListener	2017-10-26 10:36:14 -07:00
Travis Reeder	de04562b8e	Pushed triggers into start() and end()	2017-10-25 14:14:31 +02:00
Travis Reeder	d080c23981	First draft of modifying RunnerListener to CallListener to get it closer to the action (and named better).	2017-10-25 14:13:25 +02:00
Nigel Deakin	39feaf8b69	Send tracing spans to Prometheus	2017-10-20 16:30:19 +01:00
Nigel Deakin	ae31944224	Add Prometheus statistics and an example to showcase them using Grafana	2017-10-05 16:21:31 +01:00
Reed Allman	6b7b1e3c63	Merge pull request #354 from fnproject/stats Extend stats to report Failed calls	2017-09-22 10:50:59 -07:00
Nigel Deakin	54407f7b74	Extend stats to report Failed calls	2017-09-22 17:36:43 +01:00
Reed Allman	22a1b296e3	fix slot races I'd be pretty surprised if these were happening but meh, a computer running at capacity can make the runtime scheduler do all kinds of weird shit, so this locks down the behavior around slot launching. I didn't load test much as there are cries of 'wolf' running amok, and it's late, so this could be off a little -- but I think it's about this easy. cold is the only one launching slots for itself, so it should always receive its own slot (provided within time bounds). for hot we just need a way to tell the ram token allocator that we aren't there anymore, so that somebody can close the token (important). If the bug still persists then it seems likely that there is another bug around timing I'm not aware of (possible, but unlikely) or the more likely case that it's actually taking up to the timeout to launch a container / find a ram slot / find a free container. Otherwise, it's not related to the agent and the http server timeouts may need fiddling with (read / write timeout), if ruby client is failing to connect though I'm guessing that it's just that nobody is reading the body (i.e. no function runs) and the error handling isn't very well done, as we are replying with 504 if we hit a timeout (but if nobody is listening, they won't get it).	2017-09-20 10:43:12 -07:00
Nigel Deakin	ae69bb37e3	Update global stats charts to show bteakdown by function	2017-09-19 15:05:37 +01:00
Reed Allman	53ff665d69	not ready for spans yet in hot land	2017-09-08 05:06:35 -07:00
Reed Allman	4ce9163d99	nuke some TODO yey	2017-09-07 20:15:39 -07:00
Reed Allman	1811b4e230	make fn logger more reasonable something still feels off with this, but i tinkered with it for a day-ish and didn't come up with anything a whole lot better. doing a lot of the maneuvering in the caller seemed better but it was just bloating up GetCall so went back to having it basically like it was, but returning the limited underlying buffer to read from so we can ship to the db. some small changes to the LogStore interface, swapped it to take an io.Reader instead of a string for more flexibility in the future while essentially maintaining the same level of performance that we have now. i'm guessing in the not so distant future we'll ship these to some s3 like service and it would be better to stream them in than carry around a giant string anyway. also, carrying around up to 1MB buffers in memory isn't great, we may want to switch to file backed logs for calls, too. using io.Reader for logs should make #279 more reasonable if/once we move to some s3-like thing, we can stream from the log storage service direct to clients. this fixes the span being out of whack and allows the 'right' context to be used to upload logs (next to inserting the call). deletes the dbWriter we had, and we just do this in call.End now (which makes sense to me at least). removes the dupe code for making an stderr for hot / cold and simplifies the way to get a func logger (no more 7 param methods yay). closes #298	2017-09-07 20:15:39 -07:00
Reed Allman	1d0a63ca99	add id to all call invocation logs	2017-09-07 18:37:22 -07:00
Reed Allman	700078ccb9	bubble up some docker errors to user currently: * container ran out of memory (code 137) * container exited with other code != 0 * unable to pull image (auth/404) there may be others but this is a good start (the most common). notably, for both hot and cold these should bubble up (if deterministic, which hub isn't always), and these are useful for users to use in debugging why things aren't working. added tests to make sure that these behaviors are working. also changed the behavior such that when the container exits we return a 502 instead of a 503, just to be able to distinguish the fact that fn is working as expected but the container is acting funky (400 is weird here, so idk). removed references to old IsUserVisible crap and slightly changed the interface for RunResult for plumbing reasons (to get the error type, specifically). fixed an issue where if ~/.docker/config.json exists sometimes pulling images wouldn't work deterministically (should be more inline w/ expectations now) closes #275	2017-09-07 11:55:50 -07:00
Reed Allman	2341456334	FN_ prefix env vars this adds `FN_` in front of env vars that we are injecting into calls, for namespacing reasons. this will break code relying on the current variables but if we want to do this, the chance is now really. alternatively, we could maintain both the old and new for a short period of time to ease the adjustment (speak now...). updated the docs, as well. this also adds tests for the notoriously finicky configuration of the env vars and headers when setting up a call. this won't test the container / request for the call is actually receiving them, but it's a decent start and will yell loudly enough upon formatting breakage. added back FXLB_WAIT to a couple places so the lb can ride again one thing for feedback: headers are a bit confusing at the moment (not from this change, but that behavior is kept here for now), we've a chance to fix them. currently, headers in the request __are not__ prefixed with `FN_HEADER_`, i.e. 'hot'+sync containers will receive `Content-Length` in the http request headers, yet a 'cold' container from the same request would receive `FN_HEADER_Content-Length` in its environment. This is additionally confusing because if this function were hot+async, it would receive `FN_HEADER_Content-Length` in the headers, where just changing it to sync goes back to `Content-Length`. If that was confusing, then point made ;) I propose to remove the `FN_HEADER_` prefix for request headers in the environment, so that the request headers and env will match, as request headers already are of this format (not prefixed). please lmk thoughts here Would be fine with going back to the 'plain' vars too, then this patch will mostly just be adding tests and changing `FN_FORMAT` to `FORMAT`. obviously, from the examples, it's a bit ingrained now. anyway, entirely up to y'all.	2017-09-06 07:24:50 -07:00
Reed Allman	59d95d660a	push app/route cache down to datastore (#303 ) cache now implements models.Datastore by just embedding one and then changing GetApp and GetRoute to have the cache inside. this makes it really flexible for things like testing, so now the agent doesn't automagically do caching, now it must be passed a datastore that was wrapped with a cache datastore. the datastore in the server can remain separate and not use the cache still, and then now the agent when running fn 'for real' is configured with the cache baked in. this seems a lot cleaner than what we had and gets the cache out of the way and it's easier to swap in / out / extend.	2017-09-08 09:18:36 -07:00
Denis Makogon	8a337e744b	Addresing new comments	2017-09-07 15:17:39 +03:00
Denis Makogon	57a577dfc9	Wiring new context with initial span	2017-09-06 21:55:52 +03:00
Denis Makogon	9a89366d1b	Addressing review comments reverting query string caching in favour of Go 1.9 sqlx features moving context definition out of call.End to upper level	2017-09-06 21:48:29 +03:00
Denis Makogon	6a541139a9	Make call.End more solid	2017-09-06 21:48:29 +03:00
Reed Allman	4569b7fd69	stop forcing GET bodies through ?payload it's unclear why we had this behavior in the first place, but alas, no more. closes #264	2017-09-05 14:26:03 -07:00

1 2 3

102 Commits