Commit Graph
150 Commits
Author SHA1 Message Date
Alexander Kukushkin 2e62f4a2bf Merge pull request #3 from CyberDem0n/etcd-cluster-support
Etcd cluster support
2015-06-23 14:31:12 +02:00
Alexander Kukushkin b7c73fdba8 Rename hostname to name in a Member 2015-06-23 14:07:00 +02:00
Alexander Kukushkin 744026b4bb Do not update member TTL when it is far from being expired 2015-06-23 14:02:32 +02:00
Alexander Kukushkin 44e06b4cff Add expiration field into Member 2015-06-10 14:22:11 +02:00
Alexander Kukushkin 275e18e58a Refactor tests 2015-06-10 14:21:23 +02:00
Alexander Kukushkin 971bc3b6e8 Increase verbosity of failed PUT and DELETE requests. Hope will help to debug some etcd related issues 2015-06-10 11:56:16 +02:00
Alexander Kukushkin a64dad267e Bugfix: target field of SRV object should be casted to string 2015-06-09 15:46:34 +02:00
Alexander Kukushkin 57914821d0 Randomize list of etcd cluster members 2015-06-09 12:30:58 +02:00
Alexander Kukushkin a2988361c5 Retry with host from config when it failed to resolve SRV 2015-06-09 12:30:16 +02:00
Alexander Kukushkin 261fbf2721 Replace logging with logger in ha.py 2015-06-09 10:27:09 +02:00
Alexander Kukushkin da4f9e8445 Make logging of http requests more verbose 2015-06-09 10:25:33 +02:00
Alexander Kukushkin 5b7a60657f Fix syntax in .travis.yml 2015-06-09 10:04:07 +02:00
Alexander Kukushkin c79fba7656 Support work with etcd as a cluster
In case if one member of a cluster is not available it will retry with
another one and fetch the new cluster configuration. Default timeout for
all requests to etcd is 5 seconds.
Initial cluster configuration can be resolved through:
1) /v2/members call on one of the cluster members on a client port
2) when it is possible to resolve hostname into multiple ip's  it will
iterate through list and try to perform action from 1)
3) If there is discovery_srv defined in etcd section of config file it
will resolve peer addresses of all cluster members and will fetch
cluster configuration with using peer protocol by doing /members call on
a peer port
2015-06-09 09:51:53 +02:00
Alexander Kukushkin fad2fbeae0 Import ThreadingMixIn for python3 2015-06-03 16:39:29 +02:00
Alexander Kukushkin aa1fc315ae Process http requests in a threads with ThreadingMixIn 2015-06-03 16:22:05 +02:00
Alexander Kukushkin f9fbe8980f Implement lsn_to_bytes bytes_to_lsn functions 2015-06-02 13:36:36 +02:00
Alexander Kukushkin d9fbbf3ae9 Make pg_hba section more compact 2015-06-02 12:37:26 +02:00
Alexander Kukushkin fca987deda set value of api_url from application_name parameter of conn_url. Reset parameters of conn_url 2015-06-02 09:45:37 +02:00
Alexander Kukushkin f9ce29d49f rename address to conn_url in a Member obj
and introduce new field: api_url
2015-06-02 08:52:34 +02:00
Alexander Kukushkin bacd05d99e Determine preferable local address to connect through.
If listen contains '*' or 0.0.0.0 - connect via localhost
In all other cases pick the first one.
2015-06-01 17:00:15 +02:00
Alexander Kukushkin efccd777b2 rename pid_path to postmaster_pid 2015-06-01 16:47:56 +02:00
Alexander Kukushkin 1bcc2b5fa6 Extend postgres?.yml with pg_hba section to give possibility to customize pg_hba.conf 2015-06-01 16:05:35 +02:00
Alexander Kukushkin adcc7ac256 Try to avoid "double" promotion.
Also check presence of trigger_file on master after promotion when
pg_is_in_recovery() = false and if it is there - remove it.
Plus check presence of trigger_file on a new slave (after running
pg_basebackup) and if it is there - also remove it.
2015-06-01 13:48:44 +02:00
Alexander Kukushkin 8c4e547d7a Handle KeyboardInterrupt in the main loop 2015-06-01 12:02:46 +02:00
Alexander Kukushkin 939254021e Code cleanup: do not modify os.environ but pass copy of it to subprocess.call 2015-06-01 11:03:06 +02:00
Alexander Kukushkin f8c3582715 Refactor api:
Call is_running only when it's not possible to execute query against
governed postgres.
2015-06-01 10:48:16 +02:00
Alexander Kukushkin b85262637d Bugfix: do not create replication slot for master 2015-06-01 09:20:24 +02:00
Alexander Kukushkin aa486194ae Wrap time.sleep into wrapper function
In case if it was interrupted by SIGCHLD but scheduled awake time is not
reached it will continue sleeping. For all other signals behaviuor is
not changed.
2015-05-28 15:47:51 +02:00
Alexander Kukushkin cd5e589620 Revert previous commit (restore sigchld_handler) 2015-05-28 13:04:21 +02:00
Alexander Kukushkin f76f77e964 Ignore SIGCHLD instead of processing it.
This way time.sleep wont be interrupted when some external process is
finished and terminated childs would be reaped automatically.
2015-05-28 12:25:49 +02:00
Alexander Kukushkin 01e3beff1c Merge pull request #2 from CyberDem0n/restapi
Restapi
2015-05-27 18:31:42 +02:00
Alexander Kukushkin bdc4374ae3 connect address for rest api could be different from listen address 2015-05-27 14:51:41 +02:00
Alexander Kukushkin 5a817fad0d Store in etcd endpoint of restapi
In order to have backward compatibility connection url of restapi is
passed as application_name parameter in connection url of database.
Value in etcd will look like:
postgres://user:passwd@host:5432/postgres?application_name=http://host:8080/governor
Sure, this is a hack, but I guess it worth it.
2015-05-27 14:46:35 +02:00
Alexander Kukushkin 77911144c7 Improve test of governor against real config files 2015-05-27 14:44:52 +02:00
Alexander Kukushkin 56a9f52a15 Replace os.system with subprocess.call 2015-05-27 11:16:16 +02:00
Alexander Kukushkin 9968a66048 helpers/postgresql.py
Replace os.system with subprocess.call
2015-05-27 11:08:06 +02:00
Alexander Kukushkin 56daec0d6c Api: Bugfix for current_xlog_location in slaves. 2015-05-26 15:53:57 +02:00
Alexander Kukushkin 8cc0aa90b1 Fix typos when passing connection parameters to psycopg2. 2015-05-26 10:39:16 +02:00
Alexander Kukushkin 518ef68a53 Use connect_timeout=3s and statement_timeout=2s everywhere (even for local connections) 2015-05-26 09:19:20 +02:00
Alexander Kukushkin 45c4ba85b4 port for http servert must be integer 2015-05-26 09:17:58 +02:00
Alexander Kukushkin cca5bdef8f Configure listen address for restapi via a yaml file 2015-05-26 09:07:39 +02:00
Alexander Kukushkin 3e3acc3d3a Do not start listening port during test 2015-05-26 08:59:42 +02:00
Alexander Kukushkin faf295dcd0 Add initialize key into Cluster 2015-05-24 19:26:35 +02:00
Alexander Kukushkin 1096e04343 Bigfix, cursor.closed contains boolean value 2015-05-24 16:00:31 +02:00
Alexander Kukushkin db34e1f458 Write get_postgresql_status exception into log file 2015-05-24 15:57:10 +02:00
Alexander Kukushkin 755f54f266 Improve test coverage for governor.py 2015-05-24 15:42:41 +02:00
Alexander Kukushkin 8dd99170bf Refactor exceptions handling in etcd.py and ha.py
update_leader can throw EtcdError exception
unused HealthiestMemberError exception is removed
2015-05-24 15:40:36 +02:00
Alexander Kukushkin 188741de5a Merge branch 'master' of github.com:CyberDem0n/governor into restapi 2015-05-24 09:02:42 +02:00
Alexander Kukushkin ce267d9f1e query method will retry only in case of communication error 2015-05-24 09:01:23 +02:00
Alexander Kukushkin d202f72a42 Added simple web server which shows current state of postgres
Server always returns json which contains some status of postgres:
* instance type: master/slave
* xlog location for master
* xlog received_location, replayed_location
This server could be used by haproxy
GET / returns 200 if there is postgres master behind governor.
GET /slave returns 200 is there is postgres slave behind governor
In all other cases it returns 503
Currently server always listening on 0.0.0.0:8080, but it should be
configurable. In the next versions this rest api could also report some
status of etcd and could be used by is_healthiest_node method
2015-05-23 20:29:51 +02:00