Alexander Kukushkin
2e62f4a2bf
Merge pull request #3 from CyberDem0n/etcd-cluster-support
...
Etcd cluster support
2015-06-23 14:31:12 +02:00
Alexander Kukushkin
b7c73fdba8
Rename hostname to name in a Member
2015-06-23 14:07:00 +02:00
Alexander Kukushkin
744026b4bb
Do not update member TTL when it is far from being expired
2015-06-23 14:02:32 +02:00
Alexander Kukushkin
44e06b4cff
Add expiration field into Member
2015-06-10 14:22:11 +02:00
Alexander Kukushkin
275e18e58a
Refactor tests
2015-06-10 14:21:23 +02:00
Alexander Kukushkin
971bc3b6e8
Increase verbosity of failed PUT and DELETE requests. Hope will help to debug some etcd related issues
2015-06-10 11:56:16 +02:00
Alexander Kukushkin
a64dad267e
Bugfix: target field of SRV object should be casted to string
2015-06-09 15:46:34 +02:00
Alexander Kukushkin
57914821d0
Randomize list of etcd cluster members
2015-06-09 12:30:58 +02:00
Alexander Kukushkin
a2988361c5
Retry with host from config when it failed to resolve SRV
2015-06-09 12:30:16 +02:00
Alexander Kukushkin
261fbf2721
Replace logging with logger in ha.py
2015-06-09 10:27:09 +02:00
Alexander Kukushkin
da4f9e8445
Make logging of http requests more verbose
2015-06-09 10:25:33 +02:00
Alexander Kukushkin
5b7a60657f
Fix syntax in .travis.yml
2015-06-09 10:04:07 +02:00
Alexander Kukushkin
c79fba7656
Support work with etcd as a cluster
...
In case if one member of a cluster is not available it will retry with
another one and fetch the new cluster configuration. Default timeout for
all requests to etcd is 5 seconds.
Initial cluster configuration can be resolved through:
1) /v2/members call on one of the cluster members on a client port
2) when it is possible to resolve hostname into multiple ip's it will
iterate through list and try to perform action from 1)
3) If there is discovery_srv defined in etcd section of config file it
will resolve peer addresses of all cluster members and will fetch
cluster configuration with using peer protocol by doing /members call on
a peer port
2015-06-09 09:51:53 +02:00
Alexander Kukushkin
fad2fbeae0
Import ThreadingMixIn for python3
2015-06-03 16:39:29 +02:00
Alexander Kukushkin
aa1fc315ae
Process http requests in a threads with ThreadingMixIn
2015-06-03 16:22:05 +02:00
Alexander Kukushkin
f9fbe8980f
Implement lsn_to_bytes bytes_to_lsn functions
2015-06-02 13:36:36 +02:00
Alexander Kukushkin
d9fbbf3ae9
Make pg_hba section more compact
2015-06-02 12:37:26 +02:00
Alexander Kukushkin
fca987deda
set value of api_url from application_name parameter of conn_url. Reset parameters of conn_url
2015-06-02 09:45:37 +02:00
Alexander Kukushkin
f9ce29d49f
rename address to conn_url in a Member obj
...
and introduce new field: api_url
2015-06-02 08:52:34 +02:00
Alexander Kukushkin
bacd05d99e
Determine preferable local address to connect through.
...
If listen contains '*' or 0.0.0.0 - connect via localhost
In all other cases pick the first one.
2015-06-01 17:00:15 +02:00
Alexander Kukushkin
efccd777b2
rename pid_path to postmaster_pid
2015-06-01 16:47:56 +02:00
Alexander Kukushkin
1bcc2b5fa6
Extend postgres?.yml with pg_hba section to give possibility to customize pg_hba.conf
2015-06-01 16:05:35 +02:00
Alexander Kukushkin
adcc7ac256
Try to avoid "double" promotion.
...
Also check presence of trigger_file on master after promotion when
pg_is_in_recovery() = false and if it is there - remove it.
Plus check presence of trigger_file on a new slave (after running
pg_basebackup) and if it is there - also remove it.
2015-06-01 13:48:44 +02:00
Alexander Kukushkin
8c4e547d7a
Handle KeyboardInterrupt in the main loop
2015-06-01 12:02:46 +02:00
Alexander Kukushkin
939254021e
Code cleanup: do not modify os.environ but pass copy of it to subprocess.call
2015-06-01 11:03:06 +02:00
Alexander Kukushkin
f8c3582715
Refactor api:
...
Call is_running only when it's not possible to execute query against
governed postgres.
2015-06-01 10:48:16 +02:00
Alexander Kukushkin
b85262637d
Bugfix: do not create replication slot for master
2015-06-01 09:20:24 +02:00
Alexander Kukushkin
aa486194ae
Wrap time.sleep into wrapper function
...
In case if it was interrupted by SIGCHLD but scheduled awake time is not
reached it will continue sleeping. For all other signals behaviuor is
not changed.
2015-05-28 15:47:51 +02:00
Alexander Kukushkin
cd5e589620
Revert previous commit (restore sigchld_handler)
2015-05-28 13:04:21 +02:00
Alexander Kukushkin
f76f77e964
Ignore SIGCHLD instead of processing it.
...
This way time.sleep wont be interrupted when some external process is
finished and terminated childs would be reaped automatically.
2015-05-28 12:25:49 +02:00
Alexander Kukushkin
01e3beff1c
Merge pull request #2 from CyberDem0n/restapi
...
Restapi
2015-05-27 18:31:42 +02:00
Alexander Kukushkin
bdc4374ae3
connect address for rest api could be different from listen address
2015-05-27 14:51:41 +02:00
Alexander Kukushkin
5a817fad0d
Store in etcd endpoint of restapi
...
In order to have backward compatibility connection url of restapi is
passed as application_name parameter in connection url of database.
Value in etcd will look like:
postgres://user:passwd@host:5432/postgres?application_name=http://host:8080/governor
Sure, this is a hack, but I guess it worth it.
2015-05-27 14:46:35 +02:00
Alexander Kukushkin
77911144c7
Improve test of governor against real config files
2015-05-27 14:44:52 +02:00
Alexander Kukushkin
56a9f52a15
Replace os.system with subprocess.call
2015-05-27 11:16:16 +02:00
Alexander Kukushkin
9968a66048
helpers/postgresql.py
...
Replace os.system with subprocess.call
2015-05-27 11:08:06 +02:00
Alexander Kukushkin
56daec0d6c
Api: Bugfix for current_xlog_location in slaves.
2015-05-26 15:53:57 +02:00
Alexander Kukushkin
8cc0aa90b1
Fix typos when passing connection parameters to psycopg2.
2015-05-26 10:39:16 +02:00
Alexander Kukushkin
518ef68a53
Use connect_timeout=3s and statement_timeout=2s everywhere (even for local connections)
2015-05-26 09:19:20 +02:00
Alexander Kukushkin
45c4ba85b4
port for http servert must be integer
2015-05-26 09:17:58 +02:00
Alexander Kukushkin
cca5bdef8f
Configure listen address for restapi via a yaml file
2015-05-26 09:07:39 +02:00
Alexander Kukushkin
3e3acc3d3a
Do not start listening port during test
2015-05-26 08:59:42 +02:00
Alexander Kukushkin
faf295dcd0
Add initialize key into Cluster
2015-05-24 19:26:35 +02:00
Alexander Kukushkin
1096e04343
Bigfix, cursor.closed contains boolean value
2015-05-24 16:00:31 +02:00
Alexander Kukushkin
db34e1f458
Write get_postgresql_status exception into log file
2015-05-24 15:57:10 +02:00
Alexander Kukushkin
755f54f266
Improve test coverage for governor.py
2015-05-24 15:42:41 +02:00
Alexander Kukushkin
8dd99170bf
Refactor exceptions handling in etcd.py and ha.py
...
update_leader can throw EtcdError exception
unused HealthiestMemberError exception is removed
2015-05-24 15:40:36 +02:00
Alexander Kukushkin
188741de5a
Merge branch 'master' of github.com:CyberDem0n/governor into restapi
2015-05-24 09:02:42 +02:00
Alexander Kukushkin
ce267d9f1e
query method will retry only in case of communication error
2015-05-24 09:01:23 +02:00
Alexander Kukushkin
d202f72a42
Added simple web server which shows current state of postgres
...
Server always returns json which contains some status of postgres:
* instance type: master/slave
* xlog location for master
* xlog received_location, replayed_location
This server could be used by haproxy
GET / returns 200 if there is postgres master behind governor.
GET /slave returns 200 is there is postgres slave behind governor
In all other cases it returns 503
Currently server always listening on 0.0.0.0:8080, but it should be
configurable. In the next versions this rest api could also report some
status of etcd and could be used by is_healthiest_node method
2015-05-23 20:29:51 +02:00