fn-serverless

mirror of https://github.com/fnproject/fn.git synced 2022-10-28 21:29:17 +03:00

Author	SHA1	Message	Date
Tolga Ceylan	feeeca3321	fn: agent shutdown improvements (#622 )	2017-12-22 12:52:31 -08:00
Tolga Ceylan	25a72146f5	slot tracking improvements (#562 ) * fn: remove 100 msec sleep for hot containers ) moved slot management to its own file ) slots are now implemented with LIFO semantics, this is important since we do not want to round robin hot containers. Idle hot containers should timeout properly. ) each slot queue now stores a few basic stats such as avg time a call spent in a given state and number of running/launching containers, number of waiting calls in those states. ) first metrics in these basic stats are discarded to avoid initial docker pull/start spikes. ) agent now records/updates slot queue state and how much time a call stayed in that state. ) waitHotSlot() replaces the previous wait 100 msec logic where it sends a msg to hot slot go routine launchHot() and waits for a slot *) launchHot() is now a go routine for tracking containers in hot slots, it determines if a new containers is needed based on slot queue stats.	2017-12-15 15:50:07 -08:00
Tolga Ceylan	419298e1c0	Async hot hdr fix (#604 ) * fn: for async hot requests ensure/fix content-length/type * fn: added tests for FromModel for content type/length * fn: restrict the content-length fix to async in FromModel()	2017-12-15 14:32:25 -08:00
Nigel Deakin	f1fc040948	Fix spans for prometheus (#606 )	2017-12-15 10:31:57 -08:00
Tolga Ceylan	3b12f3fa3d	Fn deadline (#591 ) * fn: added fn_deadline as RFC3339	2017-12-14 19:25:36 -08:00
Tolga Ceylan	eccce881a6	fn: exclude timeouts from failed error count (#590 ) * fn: exclude timeouts from failed error count	2017-12-14 13:10:07 -08:00
Reed Allman	bb92547b95	Hybrid plumby (#585 ) * fix configuration of agent and server to be future proof and plumb in the hybrid client agent * fixes up the tests, turns off /r/ on api nodes * fix up defaults for runner nodes * shove the runner async push code down into agent land to use client * plumb up async-age * return full call from async dequeue endpoint, since we're storing a whole call in the MQ we don't need to worry about caching of app/route [for now] * fast safe shutdown of dequeue looper in runner / tidying of agent * nice errors for path not found against /r/, /v1/ or other path not found * removed some stale TODO in agent * mq backends are only loud mouths in debug mode now * update tests * Add caching to hybrid client * Fix HTTP error handling in hybrid client. The type switch was on the value rather than a pointer. * Gofmt. * Better caching with a nice caching wrapper * Remove datastore cache which is now unused * Don't need to manually wrap interface methods * Go fmt	2017-12-12 15:54:55 -08:00
Tolga Ceylan	b0937f236f	fn: headroom error case to clarify OOM (#589 )	2017-12-12 12:02:12 -08:00
Reed Allman	2ebc9c7480	hybrid mergy (#581 ) * so it begins * add clarification to /dequeue, change response to list to future proof * Specify that runner endpoints are also under /v1 * Add a flag to choose operation mode (node type). This is specified using the `FN_NODE_TYPE` environment variable. The default is the existing behaviour, where the server supports all operations (full API plus asynchronous and synchronous runners). The additional modes are: * API - the full API is available, but no functions are executed by the node. Async calls are placed into a message queue, and synchronous calls are not supported (invoking them results in an API error). * Runner - only the invocation/route API is present. Asynchronous and synchronous invocation requests are supported, but asynchronous requests are placed onto the message queue, so might be handled by another runner. * Add agent type and checks on Submit * Sketch of a factored out data access abstraction for api/runner agents * Fix tests, adding node/agent types to constructors * Add tests for full, API, and runner server modes. * Added atomic UpdateCall to datastore * adds in server side endpoints * Made ServerNodeType public because tests use it * Made ServerNodeType public because tests use it * fix test build * add hybrid runner client pretty simple go api client that covers surface area needed for hybrid, returning structs from models that the agent can use directly. not exactly sure where to put this, so put it in `/clients/hybrid` but maybe we should make `/api/runner/client` or something and shove it in there. want to get integration tests set up and use the real endpoints next and then wrap this up in the DataAccessLayer stuff. * gracefully handles errors from fn * handles backoff & retry on 500s * will add to existing spans for debuggo action * minor fixes * meh	2017-12-11 10:43:19 -08:00
Tolga Ceylan	9481f811b7	fn: fail count should include timeouts (#577 ) * fn: fail count should include timeouts	2017-12-06 16:11:59 -08:00
Nigel Deakin	96f27070be	More metrics (#561 ) * Add new spans to agent.submit * Add new spans to agent.submit * Add new spans to agent.submit * Add new spans to agent.submit	2017-12-05 10:26:28 -08:00
Travis Reeder	0798f9fac8	Middleware upgrade (#554 ) * Adds root level middleware * Added todo * Better way for extensions to be added. * Bad conflict merge?	2017-12-05 08:22:03 -08:00
Tolga Ceylan	25f6706642	Container memory tracking related changes (#541 ) * squash# This is a combination of 10 commits2 fn: get available memory related changes ) getAvailableMemory() improvements ) early fail if requested memory too large to meet ) tracking async and sync pools individually. Sync pool is reserved for sync jobs only, while async pool can be used by all jobs. ) head room estimation for available memory in Linux.	2017-12-01 11:21:16 -08:00
Travis Reeder	a67d5a6290	Drop viper dependency (#550 ) * Removed viper dependency. * removed from glide files	2017-11-28 15:46:17 -08:00
Reed Allman	892c843d87	add error to call model (#539 ) * add error to call model closes #331 previously, for async this error was being masked completely even if it was something useful like the image not existing. for sync, the error was returned in the http request but now it's also being stored. this error itself can cover a lot of landscape, it could be an error in getting a slot, pulling an image, running a container, among other things. anyway, no longer being masked. we can likely improve it in certain cases we run into in the future, but it's open ended at the moment and not being masked like some errors in sync http request returns (503 non-models.APIError) for now. * tucks in callTrigger stuff to keep api clean * adds swagger * adds migration * adds tests for datastore and agent to ensure behavior * pull images before tests are ran * gofmt migrations file	2017-11-28 11:21:39 -06:00
Nigel Deakin	954f69e74a	Add appname to basic metrics (#547 ) * Add app labels to queued/running/completed/failed metrics * Add app labels to queued/running/completed/failed metrics * Add app labels to queued/running/completed/failed metrics	2017-11-28 10:17:24 -06:00
Reed Allman	c9198b8525	add per call stats field as histogram (#528 ) * add per call stats field as histogram this will add a histogram of up to 240 data points of call data, produced every second, stored at the end of a call invocation in the db. the same metrics are also still shipped to prometheus (prometheus has the not-potentially-reduced version). for the API reference, see the updates to the swagger spec, this is just added onto the get call endpoint. this does not add any extra db calls and the field for stats in call is a json blob, which is easily modified to add / omit future fields. this is just tacked on to the call we're making to InsertCall, and expect this to add very little overhead; we are bounding the set to be relatively small, planning to clean out the db of calls periodically, functions will generally be short, and the same code used at a previous firm did not cause a notable db size increase with production workload that is worse, wrt histogram size (I checked). the code changes are really small aside from changing to strfmt.DateTime, adding a migration and implementing sql.Valuer; needed to slightly modify the swap function so that we can safely read `call.Stats` field to upload at end. with the full histogram in hand, we can compute max/min/average/median/growth rate/bernoulli distributions/whatever very easily in a UI or tooling. in particular, this data is easily chartable [for a UI], which is beneficial. * adds swagger spec of api update to calls endpoint * adds migration for call.stats field * adds call.stats field to sql queries * change swapping of hot logger to exec, so we know that call.Stats is no longer being modified after `exec` [in call.End] * throws out docker stats between function invocations in hot functions (no call to store them on, we could change this later for debug; they're in prom) * tested in tests and API closes #19 * add format of ints to swag	2017-11-27 08:52:53 -06:00
Denis Makogon	347edea56e	Use valid call type instead in protocol (#534 )	2017-11-24 10:32:17 -06:00
Tolga Ceylan	89dc79f0b0	fn: remove redundant httprouter code (#532 ) *) tree from https://github.com/julienschmidt/httprouter is already in Gin and this only seems to be parsing parameters from URI.	2017-11-22 13:58:10 -06:00
Tolga Ceylan	2551be446a	fn: introducing 503 responses for out of capacity case (#518 ) * fn: introducing 503 responses for out of capacity case ) Adding 503 with Retry-After header case if request failed during waiting for slots. ) TODO: return 503 without Retry-After if the request can never be met by this fn server. ) fn: runner test docker pull fixup ) fn: MaxMemory for routes is now a variable to allow testing and adjusting it according to fleet memory sizes.	2017-11-21 12:42:02 -08:00
Reed Allman	2d8c528b48	S3 loggyloo (#511 ) * add minio-go dep, update deps * add minio s3 client minio has an s3 compatible api and is an open source project and, notably, is not amazon, so it seems best to use their client (fwiw the aws-sdk-go is a giant hair ball of things we don't need, too). it was pretty easy and seems to work, so rolling with it. also, minio is a totally feasible option for fn installs in prod / for demos / for local. * adds 's3' package for s3 compatible log storage api, for use with storing logs from calls and retrieving them. * removes DELETE /v1/apps/:app/calls/:call/log endpoint * removes internal log deletion api * changes the GetLog API to use an io.Reader, which is a backwards step atm due to the json api for logs, I have another branch lined up to make a plain text log API and this will be much more efficient (also want to gzip) * hooked up minio to the test suite and fixed up the test suite * add how to run minio docs and point fn at it docs some notes: notably we aren't cleaning up these logs. there is a ticket already to make a Mr. Clean who wakes up periodically and nukes old stuff, so am punting any api design around some kind of TTL deletion of logs. there are a lot of options really for Mr. Clean, we can notably defer to him when apps are deleted, too, so that app deletion is fast and then Mr. Clean will just clean them up later (seems like a good option). have not tested against BMC object store, which has an s3 compatible API. but in theory it 'just works' (the reason for doing this). in any event, that's part of the service land to figure out. closes #481 closes #473 * add log not found error to minio land	2017-11-20 17:39:45 -08:00
Reed Allman	f08ea57bc0	add docker health check waiter to start (#434 ) before returning the cookie in the driver, wait for health checks https://docs.docker.com/engine/reference/builder/#healthcheck if provided. for images that don't have health checks, this will have no affect (an added call to inspect container, for hot it's small potatoes). this will be useful for containers so that they can pull large files or do setup that takes a while before accepting tasks. since this is before start, it won't run into the idle timeout. we could likely use these for hot containers in general and check between runs or something, but didn't do that here. one nascient concern is that for hot if the containers never become healthy I don't think we will ever kill them and the slot will 'leak'. this is true for this and for other cases (pulling image) I think, we should probably recycle hot containers every hour or something which would also close this. anyway, not a huge blocker I don't think, there will likely be 1 user of this feature for a bit, it's not documented since we're not sure we want to support it. closes #336	2017-11-17 20:31:33 -08:00
Tolga Ceylan	17d4271ffb	fn: move memory/token code into resource (#512 ) ) bugfix: fix nil ptr access in docker registry RoundTrip ) move async and ram token related code into resource.go	2017-11-17 15:25:53 -08:00
Nigel Deakin	910612d0b1	Docker stats to Prometheus (#486 ) * Docker stats to Prometheus * Fix compilation error in docker_test * Refactor docker driver Run function to wait for the container to have stopped before stopping the colleciton of statistics * Fix go fmt errors * Updates to sending docker stats to Prometheus * remove new test TestWritResultImpl because we changes to support multiple waiters have been removed * Update docker.Run to use channels not contextrs to shut down stats collector	2017-11-16 11:02:33 -08:00
Travis Reeder	96cfc9f5c1	Update json (#463 ) * wip * wip * Added more fields to JSON and added blank line between objects. * Update tests. * wip * Updated to represent recent discussions. * Fixed up the json test * More docs * Changed from blank line to bracket, newline, open bracket. * Blank line added back, easier for delimiting.	2017-11-16 09:59:13 -08:00
Tolga Ceylan	a530cd9be3	Minor naming and control flow changes to satisfy golint	2017-11-02 15:36:55 -07:00
Reed Allman	ce252d0448	Merge pull request #424 from fnproject/call-listener CallListener - replaces RunnerListener	2017-10-26 10:36:14 -07:00
Travis Reeder	965630af15	Remove error returns.	2017-10-26 11:12:08 +02:00
Reed Allman	91d2a89e19	Merge pull request #447 from fnproject/tracing_to_prometheus Tracing to prometheus	2017-10-25 10:58:07 -07:00
Reed Allman	7ba2dc005e	fix debug logger output (#458 ) our dear friend mr. funclogger was bypassing calls to our multi writer since we were embedding a *bytes.Buffer, it was using ReadFrom and WriteString which would never call the stderr logger's Write method (or, as I learned, other things trying to wrap that buffer's Write method...). the tl;dr is many times DEBUG lines don't get spat out, from async tasks especially (few people using this). I think the final solution is probably to make funclogger a 'more robust' interface that we understand instead of trying to minimize it to an io.ReaderWriterCloser, much like how bytes.Buffer has all kinds of methods implemented on it, we can implement things like ReadFrom and WriteString most likely. not a big fan of how things are now (and it's my own doing) with the readerwritercloser coming from multiple places but meh, will get to it some day soon, the log stuff will be a pretty hot path.	2017-10-25 16:25:59 +02:00
Travis Reeder	d30bcb0397	Fix lost error	2017-10-25 14:41:18 +02:00
Travis Reeder	de04562b8e	Pushed triggers into start() and end()	2017-10-25 14:14:31 +02:00
Travis Reeder	d080c23981	First draft of modifying RunnerListener to CallListener to get it closer to the action (and named better).	2017-10-25 14:13:25 +02:00
Nigel Deakin	39feaf8b69	Send tracing spans to Prometheus	2017-10-20 16:30:19 +01:00
Denis Makogon	ce25adfddb	JSON protocol updating (#426 ) * JSON protocol updating this patch adds HTTP query string into payload (see more TODOs in code) adds one more test to verify query * Fixing FMT	2017-10-12 23:10:21 +03:00
Nigel Deakin	1646d25c01	Merge pull request #396 from fnproject/add_prometheus_metrics Add Prometheus statistics and an example to showcase them using Grafana	2017-10-10 09:37:28 +01:00
Denis Makogon	22b5140f56	Do not expect function to set response code	2017-10-07 03:07:21 +03:00
Denis Makogon	e4684096f7	Fmt and docs	2017-10-07 02:59:08 +03:00
Denis Makogon	6141344e5f	Error before sending json object if something bad happend with reading a request body	2017-10-07 02:33:43 +03:00
Denis Makogon	6682de4768	Addressing comments	2017-10-07 02:28:56 +03:00
Denis Makogon	e8f317abd4	Addressing more comments tests do assertion on request data and headers doc fixed	2017-10-07 02:24:07 +03:00
Denis Makogon	181ccf54b4	Addressing more comments + tests	2017-10-07 02:11:49 +03:00
Denis Makogon	9f3bfa1005	Read request body and see if it's not empty then decide whether write it or not	2017-10-07 01:24:43 +03:00
Denis Makogon	b4b5302a44	Addressing certain comments from last review	2017-10-07 01:20:53 +03:00
Denis Makogon	de7b4e6067	Returning error instead of writing it to a response writer	2017-10-07 00:52:01 +03:00
Denis Makogon	7dd9b5a4cd	We still can write JSON request object in parts except just copying content from request body to STDIN we need to write encoded data, so we're using STDIN JSON stream encoder.	2017-10-07 00:43:09 +03:00
Denis Makogon	588d9e523b	Do not forget to close request body	2017-10-07 00:43:09 +03:00
Denis Makogon	c2ee67fb21	Revisiting request body processing	2017-10-07 00:43:09 +03:00
Denis Makogon	1f589d641e	Let function write headers to a response	2017-10-07 00:43:09 +03:00
Denis Makogon	caf1488dd9	Make Dispatch cleaner	2017-10-07 00:43:09 +03:00

1 2 3 4 5

246 Commits